SlipstreamJobsFresh Startup & VC-Backed Jobs

Applied ML Engineer

Sentient Foundation - Remote - Remote - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sentient Foundation is hiring an Applied ML Engineer to build systems at the intersection of machine learning research and production software. This is an end-to-end engineering role spanning model evaluation, inference infrastructure, backend systems, and product interfaces. The goal is to turn promising research methods into reliable, measurable, and usable products that users can interact with. Key responsibilities include: - Reproduce and evaluate research methods using open-weight and API-accessible models - Design evaluation datasets, probes, scoring methods, baselines, and experiment harnesses - Work directly with model weights, logits, hidden states, activations, and inference infrastructure - Build and extend evaluation infrastructure including runners, judges, persistence, and experiment orchestration - Transform research workflows into product experiences with experiment configuration, traces, comparisons, and review workflows - Investigate how verification methods behave under model modification (fine-tuning, merging, quantization, distillation, safety removal, evasion) - Design controlled experiments that isolate meaningful signals from artifacts or confounders - Write clear technical reports distinguishing measured evidence, interpretation, and hypotheses - Ship production-quality systems with APIs, background jobs, observability, testing, and documentation This role bridges research and engineering. You'll move between reading papers, building experiments, evaluating rigorously, and shipping production systems. The product surface is primarily React/TypeScript, so you should be able to make complex experiments and results understandable to users. Success in the first six months includes reproducing at least one published model-provenance or verification method, building a repeatable model-verification runner, adding verification workflows to the product UI, running controlled experiments across model variants, and improving understanding of when and why verification methods succeed or fail. This is NOT a pure research role, a generic model-training position, a frontend-only or backend-only role, or one where benchmark scores are accepted without understanding their production. The ideal candidate enjoys moving across research, experimentation, engineering, and product layers. REQUIREMENTS: - Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers - Strong understanding of ML evaluation: dataset design, baselines, metrics, calibration, false positives/negatives, statistical uncertainty, reproducibility - Ability to read ML research papers and implement methods from first principles rather than relying entirely on existing packages - Experience building production software beyond notebooks: APIs, asynchronous jobs, databases, logging, testing, deployment - Comfort working with open-weight models and understanding how modern LLM inference systems operate - Ability to work across backend and frontend boundaries - Strong technical judgment about what experimental evidence does and does not support - High agency and strong sense of ownership; comfortable identifying problems, proposing solutions, and driving work forward independently - Comfortable in fast-moving startup environments where priorities evolve and individuals operate across functions USEFUL EXPERIENCE (plus factors): - Model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, or interpretability - Activation and representation analysis, probing, model hooks, logits, hidden states, or model-internals work - Evaluation and inference infrastructure (DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, or similar) - Next.js, React, TypeScript, data visualization, or experiment dashboards - Running and serving open-weight models on GPUs; reasoning about latency, throughput, memory, precision, and cost tradeoffs - Designing adversarial evaluations or testing systems against deliberate evasion attempts

Similar roles