SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

Veeam Software - Remote - Remote - posted 2026-09-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 172,400 - 441,500 / annual

Veeam is launching a global Site Reliability Engineering (SRE) function to support the rollout and operation of Veeam Data Cloud, a new SaaS offering. As a Staff Site Reliability Engineer, you will lead by building reliable-by-default platforms, tooling, and patterns that product teams adopt at scale. You will serve as a hands-on technical leader within the SRE team, guiding senior engineers, influencing product development teams, and ensuring systems are built to be reliable, scalable, and observable from the ground up. You will drive strategic initiatives, mentor others in SRE practices, and help define architectural best practices across the platform. This role is pivotal in aligning teams, enforcing high standards, and scaling SRE principles globally. Key ownership areas include: - Building reliability features as productized code: libraries, services, and controllers (deployment safety guards, rate-limiters, circuit breakers, load-shedding adapters, back-pressure controls) that product teams import and extend. - Designing and implementing an observability platform with telemetry pipelines (metrics, logs, traces), SLI/SLOs, and error-budget policies; shipping SDKs/CLI/plugins for teams to declare SLOs in code and gate releases. - Creating a change safety toolchain with progressive delivery primitives (canary, blue/green, feature flags), automated rollback, and release validation as reusable services/operators and CI/CD integrations. - Building resilience automation: fault-injection APIs, chaos experiments, traffic shadowing, and load/perf harnesses in pre-prod and prod pipelines. - Developing golden-path platform components: Terraform/Pulumi modules, Kubernetes operators, Helm charts, and reference microservice templates with paved-road documentation. - Implementing incident learning systems with post-incident automation and code changes that remove classes of failure. Day-to-day responsibilities include writing high-quality code in Go, TypeScript/Node.js, C#, or Java; designing distributed, multi-region services (initially on Azure) with focus on failure modes and operability; partnering with Staff/Principal peers across product and platform to align on reliability standards; instrumenting systems deeply and automating detection/response; leading complex incidents and driving blameless learning; and mentoring senior engineers through design reviews, ADRs, and pair programming. Veeam's SRE philosophy emphasizes engineering over firefighting (preventing repeated incidents by changing code, not adding toil), data-driven decision-making via SLIs/SLOs and error budgets, enablement rather than gatekeeping, and blameless learning where incidents inform design. On-call expectations include standard shifts aligned to business hours in a 3×8-hour global rotation, designed for fairness with predictable schedules, protected focus time, and compensatory benefits per local policy. Requirements: - 8+ years in software engineering for cloud-based products with significant time designing and operating distributed systems at scale. - Strong proficiency in at least one backend language (C#, Java, Go, TypeScript/Node.js) and experience writing production-grade services and libraries. - Deep hands-on experience with Kubernetes, Infrastructure as Code (Terraform or Pulumi), and CI/CD platforms (GitHub Actions, GitLab, ArgoCD). - Practical observability expertise (metrics, tracing, logging) and experience turning SLOs/error budgets into engineering workflows. - Ability to lead cross-team initiatives, influence architecture, and deliver measurable reliability outcomes. - Comfort with a follow-the-sun on-call model (8-hour daytime rotations) and coverage. Nice to have: - Experience building reliability platforms (SLO/SLO policy engines, progressive delivery, chaos/validation) used by multiple teams. - Multi-cloud or advanced Azure networking/traffic management (cross-region failover, DNS, gateway, service mesh). - Performance engineering at scale (workload modeling, cost/perf tradeoffs, regression detection). - Security/compliance-aware delivery (SOC 2/ISO/SOx/FedRAMP patterns) as code.

Similar roles