SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer, Reliability (SRE)

Veeam Software - Bangalore, KA, India - In-office - posted 2026-08-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Veeam is seeking a Senior Software Engineer, Reliability to join its SRE team in Bangalore. In this hands-on technical leadership role, you will guide senior engineers, influence product development teams, and ensure systems are built to be reliable, scalable, and observable from inception. You will drive strategic SRE initiatives, mentor engineers in reliability practices, and help define architectural best practices across Veeam's global platform. Key responsibilities include designing and evolving infrastructure for high availability and fault tolerance across public clouds (primarily Azure), establishing SLIs/SLOs/error budgets, and leading incident response with blameless postmortems. You will drive adoption of deep observability practices, develop automation and self-healing tools to reduce operational toil, and participate in on-call rotations. On the engineering side, you'll contribute to infrastructure-as-code, CI/CD systems, deployment automation, and scalable configuration management. You'll integrate monitoring and chaos engineering tools to validate reliability assumptions, implement testing strategies and canary deployments, and build release validation pipelines that protect production while enabling rapid feature delivery. Collaboration is central: you'll embed within product and platform teams to champion reliability from design through delivery, mentor engineers globally, and advocate for DevOps/SRE best practices. You'll contribute to a learning culture focused on continuous improvement and proactive risk management. Required: 5+ years hands-on software engineering experience with at least 2 years in Site Reliability, Platform Engineering, or similar. Deep public cloud experience (Azure preferred), strong programming skills (JS, Node, TypeScript, Go, Java, C#, or similar), proven track record with monitoring/alerting/observability tools (Prometheus, Grafana, OpenTelemetry), experience with IaC tools (Terraform/Pulumi) and container orchestration (Kubernetes), and solid understanding of distributed systems and cloud networking.

Similar roles