SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sentient Foundation is seeking an Applied ML Engineer to bridge machine learning research and production systems. This is a full-stack engineering role requiring comfort moving across model evaluation, inference infrastructure, backend systems, and product interfaces.
You will reproduce and evaluate research methods using open-weight and API-accessible models, design rigorous evaluation datasets and scoring methods, work directly with model weights and inference infrastructure, and build evaluation systems including runners, judges, and experiment orchestration. A key responsibility is turning research workflows into product experiences—designing experiment configuration, runs, traces, comparisons, and review workflows accessible through user interfaces. You'll investigate how verification methods behave under model modification (fine-tuning, merging, quantization, distillation), design controlled experiments that isolate meaningful signals from artifacts, and write technical reports distinguishing measured evidence from interpretation.
The role requires shipping production-quality systems with APIs, background jobs, observability, testing, and documentation. You'll work across backend and frontend boundaries; the product surface is primarily React/TypeScript, and you should be able to make complex experiments understandable to users.
In the first six months, success means reproducing at least one published model-provenance or verification method with clear documentation, building a repeatable model-verification runner with versioned inputs and metrics, adding verification workflows to the product UI, running controlled experiments across model variants, and leaving behind production-quality code and tooling that peers can operate and extend.
This is NOT a pure research role, a generic model-training position, a frontend/backend-only role, or one where you accept benchmark scores without understanding their provenance. The company seeks someone who enjoys moving between research, experimentation, engineering, and product, and who cares about building trustworthy systems.
REQUIREMENTS:
- Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers
- Strong understanding of ML evaluation: dataset design, baselines, metrics, calibration, false positives/negatives, statistical uncertainty, reproducibility
- Ability to read ML research papers and implement methods from first principles rather than relying entirely on existing packages
- Experience building production software beyond notebooks: APIs, asynchronous jobs, databases, logging, testing, deployment
- Comfort with open-weight models and understanding of modern LLM inference systems
- Ability to work across backend and frontend boundaries, including React/TypeScript
- Strong technical judgment about what experimental evidence does and does not support
- High agency and ownership; comfortable identifying problems, proposing solutions, and driving work forward independently
- Ability to work in fast-moving startup environment with evolving priorities and cross-functional expectations
USEFUL EXPERIENCE (PLUS):
- Model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, or interpretability
- Activation and representation analysis, probing, model hooks, logits, hidden states, or model-internals work
- Evaluation and inference infrastructure: DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector
- Next.js, React, TypeScript, data visualization, or experiment dashboards
- Running and serving open-weight models on GPUs; reasoning about latency, throughput, memory, precision, and cost tradeoffs
- Designing adversarial evaluations or testing systems against deliberate evasion attempts