SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Veeam is seeking a Senior Software Engineer, Reliability to join its SRE team in Bangalore. You will serve as a hands-on technical leader, guiding senior engineers, influencing product development teams, and ensuring systems are built to be reliable, scalable, and observable from inception.
Key responsibilities include designing and evolving highly available, fault-tolerant infrastructure across public clouds (primarily Azure); establishing and maintaining SLIs, SLOs, and error budgets; leading incident response and blameless postmortems; driving adoption of deep observability practices with comprehensive telemetry, logs, metrics, and tracing; developing automation and self-healing tools to reduce operational toil; contributing to infrastructure-as-code, CI/CD systems, and deployment automation; integrating chaos engineering tools to validate reliability assumptions; implementing testing strategies and canary deployments; and mentoring engineers across global teams on DevOps/SRE best practices.
You will embed within product and platform teams to champion reliability from design through delivery, participate in on-call rotations, and contribute to a learning culture focused on continuous improvement and proactive risk management.
Required qualifications: 5+ years hands-on software engineering experience with at least 2 years in Site Reliability, Platform Engineering, or similar roles; deep experience building systems on public cloud providers (Azure preferred); strong programming skills in JavaScript, Node, TypeScript, Go, Java, C#, or similar languages; proven track record delivering monitoring, alerting, and observability tooling (Prometheus, Grafana, OpenTelemetry); experience with infrastructure-as-code tools (Terraform/Pulumi) and container orchestration (Kubernetes); solid understanding of distributed systems, cloud networking, and cloud-native design; excellent communication and collaboration skills across geographies.
Bonus skills include experience on large-scale B2B SaaS platforms, chaos engineering, resilience testing, performance testing, and familiarity with compliance frameworks (ISO, SOC 2, GDPR, FEDRAMP/CMMC).