SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer, Environment Automation

GitLab - Remote - Remote - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

GitLab is seeking a Staff Site Reliability Engineer specializing in Environment Automation to help keep user-facing services and production systems reliable, scalable, and efficient. This role focuses on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance—ensuring they remain secure, consistent, and reliable at scale. You will design infrastructure automation that provisions and operates GitLab environments using Terraform, Ansible, and Kubernetes. Key responsibilities include creating and maintaining deployment packages (Helm Charts, omnibus-gitlab), building and operating Dedicated GitLab instances integrated with cloud-native services (GCP, AWS), developing tools to orchestrate infrastructure-as-code workflows across multiple tenants, deploying and managing microservices on Kubernetes clusters at scale, enhancing observability stacks (Prometheus, ELK) for proactive monitoring and incident response, integrating with cloud provider ecosystems (IAM, networking, storage), and championing cloud security best practices. Day-to-day work includes building and scaling multi-tenant infrastructure using Terraform, Ansible, and Kubernetes; debugging and resolving production issues across Kubernetes clusters and cloud services; automating operations at scale through infrastructure-as-code solutions; building observability systems to detect bottlenecks and predict usage trends; leading incident response and postmortem efforts; and influencing architectural decisions around automation, scalability, and operational excellence. You will work within GitLab's Dedicated team, which is on a mission to deliver a fully managed, single-tenant GitLab experience through the GitLab Dedicated platform. The goal is to eliminate manual operations across the entire lifecycle of customer environments—provisioning, upgrades, security, and monitoring—so customers can focus on unlocking the full potential of The One DevOps Platform. REQUIREMENTS: - Proven ability to operate and troubleshoot production workloads across multiple tenants or environments; deep understanding of how distributed systems fail at scale and how to build in resilience - Strong hands-on experience with Terraform, including workspace strategies, state management, and automation patterns that scale; comfortable solving state isolation issues and building reliable, reusable infrastructure code; experience with Ansible and templating tools like Jsonnet is a plus - Skilled at diagnosing Kubernetes deployment failures, interpreting pod logs, and debugging scheduling issues and rollback scenarios in live environments; understands how pods, ReplicaSets, and controllers interact in production - Ability to read and debug code in Go and/or Ruby; familiar with identifying performance issues, scalability concerns, and contributing to infrastructure tooling through thoughtful code analysis - Experience supporting infrastructure for many customers or environments simultaneously; comfortable managing isolation, scaling, monitoring, and incident response across diverse workloads - Able to reason through complex systems and operational challenges; brings on-call experience and can lead technical discussions and incident resolution efforts under pressure - Proven ability to work across teams and with internal or external customers to solve technical problems while maintaining service commitments and clear communication - Comfortable using GitLab as a daily tool for infrastructure automation, collaboration, and operational workflows

Similar roles