SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Engineer, Code Agents Infra

Mistral - Palo Alto, CA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mistral is building full-stack AI solutions from frontier models to developer tools and compute infrastructure. This Research Engineer role focuses on designing and operating the end-to-end execution, training, and data infrastructure powering Mistral's agentic models and coding assistants. You will be a core contributor to the agent research stack, tackling critical infrastructure challenges across the agent lifecycle. Key responsibilities include: **Large-Scale Sandboxing Infrastructure**: Design and operate a high-throughput sandboxing platform executing LLM-generated untrusted code across 1M+ isolated environments concurrently for model evaluation and interactive RL environments. **Agent Data Generation Pipelines**: Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to power post-training and RL loops. **Training Codebase & Systems Optimization**: Optimize agent training codebases and distributed execution runtimes (PyTorch, Ray, SLURM/Kubernetes) to minimize multi-step rollout overhead, improve GPU utilization, and eliminate scaling bottlenecks. **Low-Latency Orchestration**: Reduce sandbox cold-start times to sub-second levels using container warm pools, snapshot/restore technology (CRIU, microVMs), and optimized image delivery layers across hybrid clusters. **Multi-Cluster Queueing & Resource Allocation**: Implement Kubernetes-native custom controllers, CRDs, and queuing systems to dynamically route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets. **Isolation & Security**: Ensure strict multi-tenant network and process isolation for untrusted agent code using container/sandboxing runtimes (gVisor, Firecracker) and default-deny network postures. **Operational Excellence**: Maintain high availability, telemetry, and automated self-healing across millions of transient jobs while participating in on-call rotations. Required qualifications: 4+ years in Systems Engineering, Distributed Systems, Cloud Infrastructure, or MLOps supporting LLM/RL workloads. Proven experience building high-throughput data processing pipelines (Ray, Spark, custom queues). Deep expertise with Kubernetes, container technologies, Linux cgroups/namespaces, and Docker optimization. Advanced proficiency in Python, Go, C++, or Rust with track record optimizing high-performance ML/backend systems. Hands-on experience with lightweight virtualization, container runtimes, or WASM. Deep familiarity with task queues, resource schedulers, and low-latency queuing for high-volume short-lived workloads. Comfort working alongside AI researchers to turn frontier ideas into production infrastructure.

Similar roles