SlipstreamJobsFresh Startup & VC-Backed Jobs

Engineering Manager

GitLab - Remote - Remote - posted 2026-09-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

GitLab is seeking an Engineering Manager to lead the globally distributed Observability team within Production Engineering. The team builds and operates the metrics, logging, alerting, and capacity planning platforms that enable GitLab engineers to understand GitLab.com and GitLab Dedicated. You will guide improvements to Prometheus-based metrics pipelines, log ingestion and retention, alerting driven by service-level objectives (SLOs), and capacity forecasting. In this role, you will lead, hire, onboard, and develop a distributed engineering team working asynchronously across time zones. You'll set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams, and help the team deliver observability services iteratively. You own the reliability, scalability, and cost of the team's metrics, logging, alerting, and capacity planning platforms. Key responsibilities include reducing noisy or missing alerts and telemetry gaps, using SLOs and error budgets to help engineers maintain service health, and guiding technical decisions about time-series storage, high-cardinality metrics, log pipelines, and distributed tracing. You'll participate in the Incident Manager On Call (IMOC) rotation, coordinating responses to high-severity incidents. You'll also keep the team's on-call rotation sustainable through coverage across time zones, useful runbooks, better alerts, and follow-through on post-incident actions. Additionally, you'll use AI tools and agents to support engineering workflows and incident triage, reviewing their output while engineers retain responsibility for decisions. GitLab embraces AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. The company operates with high-performance culture driven by values and continuous knowledge exchange, enabling team members to reach their full potential while collaborating with industry leaders. **Requirements:** - Experience leading an observability, platform engineering, or site reliability engineering team operating at scale, including supporting people in a distributed, asynchronous environment - Technical knowledge of metrics systems such as Prometheus and long-term storage, logging platforms such as Elasticsearch or cloud-native services, and alerting design - Experience using SLOs, error budgets, and capacity forecasts to make reliability and investment decisions - Experience operating a large software-as-a-service platform and investigating production issues such as telemetry gaps, ingestion limits, or noisy and missing alerts - Experience participating in and improving production on-call rotations, including incident coordination and balancing operational load with project work - Ability to explain technical tradeoffs to engineering partners and other stakeholders - Experience using AI tools or agents in engineering or management work; ability to describe how you would apply them to operational problems such as incident triage GitLab welcomes different paths into this role, whether through practical experience, formal study, or transferable skills.

Similar roles