SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 220,000 - 265,000 / annual
Legion is seeking a Director of Production Engineering to lead DevOps, SRE, and security operations teams responsible for the availability, scalability, and security of their production environment. This is a hands-on leadership role where you'll spend 20-30% of your time contributing directly to architecture, tooling, and incident response, with the remainder focused on vision, roadmap, and cross-team execution.
Key responsibilities include hiring and building a globally-distributed DevOps/SRE engineering team; owning the reliability and infrastructure roadmap for AWS-based production (EKS, RDS, and related services); leading the organization's security operations practice including vulnerability management, threat detection, and incident response; defining and driving engineering OKRs for infrastructure reliability, automation, and security; championing observability and alerting best practices using tools like Datadog; driving Infrastructure-as-Code, CI/CD, and automation practices; ensuring platform security, compliance, and data protection; and leading incident management on-call rotations.
Required qualifications include 8-12 years of DevOps, SRE, or production infrastructure experience with people management; deep hands-on AWS experience (EKS, RDS, VPC, IAM, Lambda, S3); demonstrated security operations experience; 5+ years with observability platforms (Datadog, Prometheus, Grafana); strong Infrastructure-as-Code experience (Terraform, CloudFormation); proficiency in Go, Python, or Bash; hands-on Linux/Unix production experience; proven cross-functional partnership skills; and a Bachelor's degree in Computer Science or related field (Master's preferred).
Preferred qualifications include experience with other cloud providers (GCP, OCI), security certifications (CISSP, AWS Security Specialty, CKS), compliance framework experience (SOC 2, ISO 27001, HIPAA), 3+ years Kubernetes experience at scale, Kubernetes-native tooling experience (Argo Workflows, Helm), 5+ years leading teams in agile environments, and experience building automated investigation and remediation pipelines.