SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Veeam is seeking a Senior Platform Engineer to join the Workload team within R&D. You will own critical observability infrastructure, drive incident response maturity, and help scale proactive support capabilities for Veeam's BaaS platform.
Key responsibilities include:
- Design, build, and maintain observability pipelines using the Elastic Stack (Elasticsearch, Kibana, Fleet) across Azure and AWS workloads
- Develop and own SLO/SLI dashboards and error budget reporting for BaaS platform services
- Lead incident response for distributed, multi-tenant cloud workloads; create and maintain runbooks
- Build proactive support tooling including pattern analysis, tenant correlation dashboards, and baseline deviation alerting
- Manage Elastic Fleet agent policies, enrollment health, and log streaming pipelines across Azure and AWS worker fleets
- Partner with SRE, R&D, and Proactive Support teams to close observability gaps and integrate with admin portals
Required experience:
- 5+ years in cloud platform engineering, SRE, or infrastructure roles supporting commercial SaaS
- Deep hands-on experience with Elastic Stack: dashboards, KQL/Query DSL, Fleet management
- Proven experience operating distributed, multi-tenant workloads on Azure and/or AWS
- Strong Azure knowledge: AKS, Entra ID, Key Vault, Service Bus, Cosmos DB, Private Endpoints
- Production incident response experience with runbook development and post-incident review
- Infrastructure as Code (Azure Bicep, Terraform) and CI/CD pipelines (Azure DevOps, GitHub Actions)
- Strong scripting in Bash, Python, or PowerShell
- Cross-functional collaboration with SRE, product, and support teams
Bonus: familiarity with Veeam Data Platform products.