SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 223,200 - 380,400 / annual
GitLab is seeking a Principal Site Reliability Engineer to lead platform and operating-model transformation for GitLab Dedicated, a fully managed single-tenant SaaS offering. This is a highly influential technical leadership role that will shape the next phase of scaling a growing fleet of isolated, customer-specific environments while maintaining reliability, security, and compliance.
In this role, you will set technical direction for GitLab Dedicated's architecture and platform strategy. You'll lead platform transformations across resilience and failover, tenant orchestration, change management, self-service tooling, and platform integrations. A key focus is driving scalable, modular architecture that aligns Dedicated with GitLab's broader Cells strategy while preserving security, isolation, and compliance requirements.
You will strengthen service ownership and operational maturity across engineering teams, helping them build, operate, and improve the production systems they own. Using production signals and incident patterns, you'll identify and address systemic reliability and scalability risks. You'll establish reusable platform patterns and automation to reduce operational toil and enable efficient scaling.
As a Principal Engineer, you'll lead complex technical decisions across teams, balancing reliability, security, cost, maintainability, and customer needs. You'll advance engineering excellence through architectural leadership, mentorship of senior engineers, and influence across the organization.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and over 50% of the Fortune 100 trust GitLab.
REQUIREMENTS:
- Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering, with experience designing and operating large-scale production systems
- Hands-on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices
- Strong software engineering fundamentals, with experience building production systems or infrastructure tooling in languages such as Go, Ruby, Python, or similar
- Strong distributed systems and systems-design expertise, with sound judgment around reliability, failure isolation, scalability, and operational complexity
- Track record of technical leadership across multiple teams, setting direction and driving complex initiatives through influence
- Experience leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems through major growth
- Experience leading changes that improve how engineering teams own and operate production systems, strengthening reliability, operational readiness, and accountability at scale
- Exceptional technical communication and influence, with ability to build alignment, mentor senior engineers, and guide complex architectural decisions