SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

Vannevar Labs - San Diego, CA, United States - Hybrid - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Vannevar Labs is a defense technology company building AI systems to deter adversaries and model adversary behavior for decision-makers. The company has grown from $3M to $80M ARR in three years and achieved unicorn status. You will own the reliability, health, and deployment automation of Vannevar's platform. As the Site Reliability Engineer, you'll monitor system dashboards and telemetry to detect health issues and performance degradation before they become incidents. You'll own the end-to-end debugging and incident response process, exercising judgment about when to investigate deeply versus escalate to the right team members. Key responsibilities include: - Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks - Own debugging and incident response processes end-to-end - Build logging, monitoring, and observability tooling to visualize platform state and mature SRE practices - Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning - Understand and improve the deployment process; automate build and deployment pipelines - Identify bottlenecks in engineering workflows and drive improvements - Develop self-service tools and automation to improve engineering efficiency - Implement secure, robust, high-availability delivery pipelines - Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders The role emphasizes simple systems that are easier to understand, maintain, and scale. You'll work in a small, agile team combining world-class engineers with veteran strategists. Clear, calm communication during incidents and in day-to-day work is essential. REQUIREMENTS: - 5+ years of experience in SRE, DevOps, or software engineering - Hands-on experience monitoring production systems and responding to incidents; comfortable owning a debugging process and making the call on when to dig in versus escalate - Excellent communication skills, especially the ability to stay clear and organized while troubleshooting live issues - Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools - Experience participating in an on-call rotation and running or contributing to post-mortems - Knowledge of AWS cloud technologies - Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi - Experience with Python, Bash, or other scripting languages - Experience working in an agile scrum environment, with the ability to work independently - Able to quickly learn new and existing technologies - Strong attention to detail and analytical capabilities - Willingness and ability to work on-site in San Diego, CA - U.S. Citizenship status required (access to U.S.-only data systems and export-controlled data) - TS/SCI Clearance required NICE-TO-HAVES: - Experience defining and tracking SLOs/SLIs and error budgets - Experience crafting CI/CD processes and automation - Proficient with containerization technologies like Docker - Experience working in AWS GovCloud - Experience with modern web services architectures - Experience with relational database systems, including SQL and relational design - Experience working with Elasticsearch/OpenSearch - Strong collaboration and negotiation skills for cross-functional projects

Similar roles