SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Grafana Labs is seeking a Staff Software Engineer to join the Application Core Services (AppCore) team, which builds the control plane powering Grafana Cloud. AppCore owns the distributed systems responsible for creating, configuring, reconciling, and operating thousands of customer environments safely and reliably.
In this role, you will design and build large-scale backend systems that maintain customer stack state consistency across Grafana Cloud. Key responsibilities include developing reconciliation and automation workflows to improve reliability and reduce operational complexity, leading technical initiatives spanning multiple services in collaboration with Product, Infrastructure, and adjacent platform teams, and improving deployment workflows, regional expansion, and incident recovery across customer environments.
You will own the production systems you build, including observability, operational tooling, on-call participation in a global follow-the-sun rotation, and continuous reliability improvements. The role involves working at the intersection of product, platform, and operations, solving complex distributed systems challenges around reliability, automation, and operational efficiency.
Grafana Labs is a 100% remote, open-source-first company with 1,600+ team members across 40+ countries. The company is backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, and others. You'll have access to modern AI coding assistants (OpenAI, Anthropic, Google frontier models) with company-funded usage budgets to accelerate development, prototyping, testing, and documentation.
The ideal candidate has experience building and operating large-scale SaaS or cloud platforms, strong backend engineering expertise in Go or similar systems languages, hands-on experience with Kubernetes and infrastructure-as-code, and enjoys solving distributed systems problems around eventual consistency and system reliability. You should be comfortable independently leading projects, navigating ambiguity, and contributing to both internal systems and the broader open-source community.