SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Veeam is building a global SRE function to support Veeam Data Cloud, a new SaaS platform. This role focuses on the Government and Sovereign Cloud environment, which operates with restricted access to government infrastructure. You'll be part of a small, high-ownership team responsible for the full platform stack, including all VDC workloads. You won't hand off problems to other teams; you'll need to understand the entire architecture deeply and own it end-to-end. You'll get up to speed on the platform quickly, often by reading code, documentation, and architecture artifacts rather than having direct environment access from day one.
This is a ground-up role where you'll help define how reliability engineering operates here. Key responsibilities include: discovering and documenting the full platform, all VDC workloads, dependencies, and risk areas; working with subject matter experts to fill knowledge gaps and build onboarding materials; writing and maintaining runbooks, architecture docs, and operational guides; designing infrastructure for high availability and fault tolerance on Azure (including Azure Government); defining SLIs, SLOs, and error budgets; running incident response and blameless postmortems; identifying reliability risks across modern and legacy workloads and building practical remediation plans within compliance constraints; closing observability gaps by defining instrumentation requirements and driving implementation; setting alerting, telemetry, and monitoring standards; building automation to reduce toil and support fleet management; participating in on-call rotations; working with IaC, CI/CD, deployment automation, and config management in air-gapped or compliance-restricted environments; building and maintaining testing, canary deployment, and release validation pipelines; and integrating chaos engineering and monitoring tools adapted to meet regulatory requirements.