SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Delinea is seeking an experienced Site Reliability Engineer to own the availability, performance, and reliability of critical SaaS applications running on Azure across multiple geographic regions. You will drive automation, monitoring, incident response, and infrastructure improvements in a multi-cloud, multi-region environment, working closely with senior engineering and cross-functional teams.
Key responsibilities include:
- Owning production SaaS application availability and performance on Azure (AKS, App Service, Redis, SQL, Service Bus) across multiple regions
- Leading troubleshooting and resolution of cloud infrastructure and application issues, including Kubernetes pod/node failures, deployment rollbacks, networking, and autoscaling problems
- Participating in on-call rotation (including weekends) and driving incident response from detection through resolution with focus on customer impact minimization
- Driving improvements to disaster recovery, failover, and incident management processes across multi-region deployments
- Building and maintaining automation scripts and monitoring tools to reduce manual toil
- Authoring post-incident reviews (RCAs), identifying root causes, and driving preventive actions to closure
- Partnering with senior engineers to implement reliability, observability, and performance best practices
- Contributing to continuous improvement initiatives across infrastructure, tooling, and process
- Communicating clearly with customer-facing stakeholders during incidents requiring external status updates
Required qualifications:
- 5+ years of Site Reliability Engineering, DevOps, or Cloud Administration experience with demonstrated production system ownership
- Hands-on Azure administration experience including AKS (Kubernetes), core Azure services, cloud networking, and security fundamentals
- Solid understanding of monitoring, logging, and alerting practices (Datadog, Azure Monitor, ELK stack) with hands-on troubleshooting using log analysis and APM tools
- Networking fundamentals knowledge: firewalls, load balancers, VPNs, DNS, routing
- Automation and scripting experience (PowerShell, Python, or similar)
- Practical understanding of backup, redundancy, and disaster recovery strategies in cloud environments including geo-redundant and multi-region deployments
- Strong ownership mindset across the full incident lifecycle with customer-first approach
- Clear and professional written communication skills for incident updates and post-incident summaries
Desired experience includes AWS Cloud Platform, CI/CD tools (Azure DevOps), infrastructure-as-code tools (Terraform, ARM templates), and operating SaaS products with regional tenant architectures.