SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Snowflake's Frontier Security AI team is building production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers. As a Staff Software Engineer, you will lead the design and development of the Agentic Harness—a runtime platform that provides the execution environment, tools, context, state, policies, and observability required to build and operate production AI agents—and an agent evaluation platform that measures task completion and detects regressions before production deployment.
This is a hands-on technical leadership role combining architecture, implementation, and mentorship. You will architect and build the Agentic Harness for executing complex multi-step AI workflows across models, tools, data, and services. You'll design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. You own agent quality end-to-end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates.
Key responsibilities include converting ambiguous quality reports into measurable failure modes and reproducible tests; analyzing production agent trajectories to identify failures in reasoning, retrieval, tool use, context, orchestration, and application code; developing offline and online measurements for task completion, correctness, groundedness, safety, latency, reliability, and cost; building simulation and replay infrastructure for golden-set tests, adversarial scenarios, model comparisons, and large-scale experiments; improving agent efficiency through model routing, prompt and semantic caching, context compaction, and token optimization; productionizing new model capabilities as secure, observable, multi-tenant services; establishing standards for evaluation design including sampling, ground-truth quality, grader calibration, and statistical significance; defining technical direction across multiple teams; and mentoring engineers while remaining directly involved in implementation and debugging.
You will work across product, infrastructure, applied AI, security, and modeling teams to take new AI capabilities from prototype to dependable customer value, closing the loop between production incidents, root-cause analysis, evaluation coverage, and regression prevention.
REQUIREMENTS:
• 9+ years of software engineering experience, including technical leadership of complex production systems
• Direct experience shipping and operating LLM applications, AI agents, or model-backed workflows in production
• Strong background in distributed systems, service architecture, high-throughput APIs, concurrency, and failure handling
• Experience building an agent runtime, workflow engine, developer platform, evaluation system, or similar infrastructure
• Demonstrated ability to evaluate nondeterministic systems without relying on a single aggregate score
• Fluency in Python and strong proficiency in at least one systems or application language (Java, Go, Rust, or TypeScript)
• Hands-on knowledge of tool calling, structured generation, retrieval, context engineering, prompt management, and model APIs
• Experience with production observability, including structured traces, replay, metrics, logs, and incident diagnosis
• Ability to balance agent quality with latency, reliability, security, and inference cost
• Track record of setting technical direction and delivering results across organizational boundaries
• Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
• Clear written and verbal communication with engineering, product, and leadership audiences
BONUS EXPERIENCE:
• Building evaluation or observability infrastructure for agentic coding, data engineering, or analytics systems
• Designing human-evaluation programs, scoring rubrics, annotation workflows, or grader-calibration methods
• Working with multi-agent orchestration, long-running agents, asynchronous workflows, or durable execution
• Developing synthetic tasks, simulations, adversarial tests, red-team exercises, or safety guardrails
• Building retrieval systems using vector search, hybrid search, semantic indexing, ranking, or caching
• Operating multi-tenant systems processing sensitive enterprise data
• Working with model training, fine-tuning, reinforcement learning, or feedback-driven optimization
• Evaluating and onboarding frontier models based on measured product outcomes
• Experience with databases, SQL engines, data platforms, Kubernetes, or cloud-native infrastructure