SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Snorkel AI is hiring Senior and Staff AI Engineers to build the infrastructure powering AI development at scale. The company's platform enables teams to create, experiment with, evaluate, and operate LLM and agentic workloads—from synthetic data generation and evaluation pipelines to simulation environments, orchestration systems, and LLM infrastructure.
You'll work at the intersection of distributed systems and applied AI, designing systems that make non-deterministic AI workloads observable, reproducible, measurable, and scalable. Key responsibilities include:
- Design and build infrastructure for large-scale agentic workloads, including multi-step agents interacting with tools, external services, sandboxes, and simulated environments
- Build scalable synthetic data generation and automated labeling systems for creating and evaluating high-quality training datasets
- Design evaluation infrastructure for measuring AI system behavior across models, prompts, tools, and multi-step trajectories, including reproducible experiments and regression detection
- Build orchestration and distributed compute systems for running thousands to millions of AI experiments reliably across heterogeneous environments
- Develop infrastructure for agent simulation environments, including provisioning, isolation, and lifecycle management
- Build and operate LLM infrastructure for routing, rate limiting, caching, provider failover, cost attribution, and efficient multi-provider execution
- Instrument workloads for observability—capturing traces, model interactions, tool calls, environment state, and cost metrics
- Design systems that make non-deterministic workloads reproducible, enabling engineers to compare experiments and diagnose behavioral regressions
- Improve developer experience through APIs, SDKs, workflow abstractions, and tooling for moving workloads from local development to production
- Collaborate with research, product, and engineering teams to turn experimental workflows into reliable platform capabilities
Required: 5+ years building production software systems with experience in AI/ML infrastructure, ML platforms, distributed systems, or backend infrastructure. You must have hands-on experience operating non-deterministic AI/ML workloads in production at significant scale. Strong proficiency in Python, production-quality API design, and distributed systems fundamentals (AWS preferred). Experience with experimentation infrastructure, model development, synthetic data, agentic workflows, or production ML systems is essential.