SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff - Engineering

Patronus AI - San Francisco, CA, United States - In-office - posted 2026-09-19

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 175,000 - 300,000 / annual

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. The company is behind influential AI evaluation research including FinanceBench, Lynx, SimpleSafetyTests, CopyrightCatcher, and Humanity's Last Exam. The team includes former researchers and engineers from Meta AI, Amazon AGI, and Google, and is backed by investors including Lightspeed Venture Partners, Notable Capital, Stanford University, and notable individuals like Noam Brown and Gokul Rajaram. Customers include foundation model labs and Fortune 500 enterprises like Adobe. As a Member of Technical Staff – Engineering, you will build the systems, infrastructure, and products powering simulation research and agent training work. This is a broad engineering role for people who operate across boundaries. Depending on the problem, you might build a realistic RL environment end-to-end, design infrastructure for running thousands of agent trajectories, ship internal platforms used by researchers, deploy and serve models, or build AI-powered developer tools. Key responsibilities include: - Building agent environments and simulations end-to-end, including frontend interfaces, backend services, APIs, data models, tools, and realistic workflows for training and evaluating AI agents - Building infrastructure powering the agent gym, including orchestration, sandboxing, packaging, benchmarking, and systems for running environments across heterogeneous targets - Developing internal platforms and developer tools used by researchers and engineers, from backends and dashboards to CLIs, SDKs, review agents, codegen helpers, and workflow automations - Building and operating ML infrastructure, including model deployment and serving, evaluation systems, GPU workloads, and services making compute accessible to the team - Owning systems from ambiguous idea through production: defining problems, making architectural decisions, implementing solutions, instrumenting them, and iterating based on real-world performance - Designing for correctness, edge cases, adversarial agent behavior, reproducibility, observability, and production system realities - Partnering closely with researchers to productionize experiments and build software and infrastructure for scalable systems - Using AI coding tools like Claude Code, Codex, and Cursor fluently, and building new tooling and automations when existing workflows are insufficient - Moving quickly without sacrificing judgment, making pragmatic decisions about what needs robustness today versus later evolution You will work across the full stack—frontend interfaces, backend services, infrastructure, and ML/agent integrations—partnering closely with researchers and engineers to turn ambiguous problems into robust systems. REQUIREMENTS: - Track record of shipping non-trivial software end-to-end as an individual contributor, ideally at a startup or on a small, high-velocity team - Strong engineering fundamentals and meaningful depth in at least one of: backend/infrastructure, frontend/product engineering, or ML systems, with ability and desire to work across boundaries - Production system experience in languages such as Python, Go, and/or TypeScript, with ability to become productive quickly in unfamiliar stacks - Experience working with modern LLMs and agents at the application level, including tool calling, agent loops, context management, harnesses, and evaluation - Strong engineering judgment around system design, correctness, reliability, failure modes, edge cases, and operational complexity - High independence: ability to take ambiguous goals, determine what needs building, find necessary people or information to unblock yourself, and ship without fully pre-scoped work - Fluency with modern AI coding tools and strong instinct for where AI can automate or accelerate engineering workflows - BS, MS, or PhD in Computer Science, Machine Learning, Software Engineering, or related quantitative field—or equivalent experience Desirable experience (depending on area of depth): - Building complex full-stack products using React, TypeScript, Next.js, Python, relational databases, and modern API frameworks - Building developer platforms, distributed systems, orchestration systems, sandboxes, or internal infrastructure - Reinforcement learning environments, agent evaluation, verifiers, reward models, or benchmarking infrastructure - Deploying and serving ML models using managed inference providers or self-operated GPUs - GPU infrastructure, workload schedulers, Kubernetes, containers, CI/CD, and cloud infrastructure - MLOps/LLMOps tooling, experiment tracking, model registries, and production observability - Browser automation tools such as Playwright or Selenium - Building or deeply using complex enterprise software and understanding operational edge cases in real-world systems

Similar roles