SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, Site Reliability Engineering

DuckDuckGo - Remote - Remote - posted 2026-09-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 243,800 - 243,800 / annual

DuckDuckGo is a profitable, remote-first online protection company founded in 2008 with 300+ team members and annual revenue exceeding $100M USD. The company operates a privacy-focused search engine, browser (Mac, Windows, iOS, Android), and Duck.ai—a private AI chat platform integrating ChatGPT, Claude, and other models. As Director of Site Reliability Engineering, you will lead the Site Reliability team in building and maintaining world-class infrastructure serving millions of users. You'll oversee complex operational challenges spanning software, systems, automation, and process analysis. Recent team projects include ensuring Duck.ai meets reliability standards, scaling the company's own document index infrastructure to handle billions of documents, and implementing privacy-respecting anti-fraud verification systems. You'll work with high-level languages including Perl, Go, TypeScript, and Python. The role requires deep technical expertise: you'll read, write, troubleshoot, and deploy software across large-scale deployments; lead and collaborate on high-impact projects from proposal through postmortem; root-cause instability in distributed systems; and partner closely with software engineers on production triage and remediation. You'll also design and implement agentic AI workflows, leverage cloud-native services and Docker-based deployments, and help chart the technical direction of the company's infrastructure. DuckDuckGo emphasizes end-to-end ownership, trust, inclusivity, and empowered project management. The company is remote-first with flexible work arrangements (no core hours, ~40 hours/week expected). All team members are required to attend an all-hands meetup and team retreat annually (each 4–5 days). Video conferencing is required for meetings. Background check required. REQUIREMENTS: - 10+ years relevant professional experience in reliability, platform, infrastructure, or software engineering - 4+ years leading SRE teams - Experience participating in 24x7 on-call rotation for large-scale deployment - Proficiency in AI-driven development, including designing and implementing agentic workflows - Deep experience administering and troubleshooting Linux and web technologies - Advanced programming skills enabling close partnership with software engineers on production issues - Hands-on experience with Docker and Docker Compose for application packaging and deployment - Ability to lead complex projects, wrangle vague problems, propose innovative solutions, and execute with strong focus on metrics - Experience developing effective tools, services, alerts, and responses to identify and address reliability risks - Ability to implement automation around infrastructure provisioning and configuration management - Foresight to identify future technical direction of deployment for improved reliability and performance - Ability to leverage cloud-native services and architectures

Similar roles