SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Frontier Security Team

Snowflake - Menlo Park, CA, United States - In-office - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Snowflake is seeking a Staff Software Engineer to lead the design and development of the Agentic Harness and agent evaluation platform for the Frontier Security AI team. This role sits at the intersection of production AI systems, infrastructure, and security, requiring deep technical leadership across multiple teams. You will architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. Key responsibilities include designing stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. You will own agent quality end-to-end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. A critical part of this role is converting ambiguous production issues (e.g., "the agent feels worse") into measurable failure modes, reproducible tests, and durable fixes. You will analyze production agent trajectories to identify failures in reasoning, retrieval, tool use, context, orchestration, and application code, then close the loop between incidents, root-cause analysis, evaluation coverage, and regression prevention. You will develop offline and online measurements for task completion, correctness, groundedness, safety, latency, reliability, and cost. You will build simulation and replay infrastructure for golden-set tests, adversarial scenarios, model comparisons, and large-scale experiments. You will improve agent efficiency through model routing, prompt and semantic caching, context compaction, tool-result management, and token optimization. Additional responsibilities include productionizing new model capabilities as secure, observable, multi-tenant services with clear operational controls, establishing standards for evaluation design (sampling, ground-truth quality, grader calibration, leakage prevention, statistical significance), defining technical direction across multiple teams, mentoring engineers, and remaining directly involved in implementation and debugging. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers—products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. REQUIREMENTS: - 9+ years of software engineering experience, including technical leadership of complex production systems - Direct experience shipping and operating LLM applications, AI agents, or model-backed workflows in production - Strong background in distributed systems, service architecture, high-throughput APIs, concurrency, and failure handling - Experience building an agent runtime, workflow engine, developer platform, evaluation system, or similar infrastructure - Demonstrated ability to evaluate nondeterministic systems without relying on a single aggregate score - Fluency in Python and strong proficiency in at least one systems or application language (Java, Go, Rust, or TypeScript) - Hands-on knowledge of tool calling, structured generation, retrieval, context engineering, prompt management, and model APIs - Experience with production observability, including structured traces, replay, metrics, logs, and incident diagnosis - Ability to balance agent quality with latency, reliability, security, and inference cost - Track record of setting technical direction and delivering results across organizational boundaries - Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience - Clear written and verbal communication with engineering, product, and leadership audiences BONUS EXPERIENCE: - Building evaluation or observability infrastructure for agentic coding, data engineering, or analytics systems - Designing human-evaluation programs, scoring rubrics, annotation workflows, or grader-calibration methods - Working with multi-agent orchestration, long-running agents, asynchronous workflows, or durable execution - Developing synthetic tasks, simulations, adversarial tests, red-team exercises, or safety guardrails - Building retrieval systems using vector search, hybrid search, semantic indexing, ranking, or caching - Operating multi-tenant systems processing sensitive enterprise data - Working with model training, fine-tuning, reinforcement learning, or feedback-driven optimization - Evaluating and onboarding frontier models based on measured product outcomes - Experience with databases, SQL engines, data platforms, Kubernetes, or cloud-native infrastructure

Similar roles