SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Frontier Security Team

Snowflake - Bengaluru, India - In-office - posted 2026-09-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Snowflake is hiring a Staff Software Engineer for the Frontier Security AI team to lead the design and development of production-grade LLM applications, intelligent agents, and AI infrastructure for enterprise customers. This is a technical leadership role focused on building the Agentic Harness—a system that executes complex, multi-step AI workflows—and an agent evaluation platform that ensures quality, security, reliability, and efficiency at scale. Key responsibilities include: - Architect and build the Agentic Harness for executing multi-step AI workflows across models, tools, data, and services, with stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. - Own agent quality end-to-end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. - Convert ambiguous quality reports into measurable failure modes, reproducible tests, and durable fixes. - Analyze production agent trajectories to identify failures in reasoning, retrieval, tool use, context, orchestration, and application code. - Close the loop between production incidents, root-cause analysis, evaluation coverage, and regression prevention. - Develop offline and online measurements for task completion, correctness, groundedness, safety, latency, reliability, and cost. - Build simulation and replay infrastructure for golden-set tests, adversarial scenarios, model comparisons, and large-scale experiments. - Improve agent efficiency through model routing, prompt and semantic caching, context compaction, tool-result management, and token optimization. - Productionize new model capabilities as secure, observable, multi-tenant services with clear operational controls. - Establish standards for evaluation design, including sampling, ground-truth quality, grader calibration, leakage prevention, and statistical significance. - Define technical direction across multiple teams and lead projects whose scope extends beyond a single service. - Mentor engineers, raise the quality of architecture reviews, and remain directly involved in implementation and debugging. Requirements: - 9+ years of software engineering experience, including technical leadership of complex production systems. - Direct experience shipping and operating LLM applications, AI agents, or model-backed workflows in production. - Strong background in distributed systems, service architecture, high-throughput APIs, concurrency, and failure handling. - Experience building an agent runtime, workflow engine, developer platform, evaluation system, or similar infrastructure. - Demonstrated ability to evaluate nondeterministic systems without relying on a single aggregate score. - Fluency in Python and strong proficiency in at least one systems or application language (Java, Go, Rust, or TypeScript). - Hands-on knowledge of tool calling, structured generation, retrieval, context engineering, prompt management, and model APIs. - Experience with production observability, including structured traces, replay, metrics, logs, and incident diagnosis. - Ability to balance agent quality with latency, reliability, security, and inference cost. - Track record of setting technical direction and delivering results across organizational boundaries. - Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. - Clear written and verbal communication with engineering, product, and leadership audiences. Bonus qualifications include: building evaluation or observability infrastructure for agentic coding, data engineering, or analytics systems; designing human-evaluation programs and grader-calibration methods; experience with multi-agent orchestration, long-running agents, or durable execution; developing synthetic tasks, simulations, adversarial tests, or safety guardrails; building retrieval systems with vector search, hybrid search, or semantic indexing; operating multi-tenant systems processing sensitive enterprise data; experience with model training, fine-tuning, or reinforcement learning; evaluating and onboarding frontier models; and experience with databases, SQL engines, data platforms, Kubernetes, or cloud-native infrastructure.

Similar roles