SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Delinea is seeking a Senior Site Reliability Engineer to join the Cloud Engineering team and own the availability and performance of production SaaS services running in Azure and AWS, including a FedRAMP High environment. This is a hands-on role focused on measuring reliability, tuning detection, responding to incidents, and automating manual work.
Key responsibilities include:
- Own reliability end-to-end for a set of production SaaS services: availability, performance, and capacity planning
- Define and manage service-level indicators (SLIs), service-level objectives (SLOs), and error budgets to prioritize reliability work
- Build and tune monitoring in Datadog and Azure Monitor, including threshold, composite, and anomaly detection monitors, synthetic checks, dashboards, and alert routing
- Automate incident response by wiring monitors into remediation workflows and replacing manual runbook steps with code
- Participate in on-call rotation, lead incident response for high-severity events, and coordinate resolution across teams
- Write post-incident reviews and customer-facing root cause analyses, then drive preventive actions to completion
- Work support escalations: reproduce issues, diagnose from logs/traces/network captures, and feed patterns back to product
- Build and maintain infrastructure as code using Terraform and Azure DevOps pipelines
- Administer the web application firewall: rule tuning, rate limiting, false positive triage, and security team coordination
- Own observability platform costs: manage Datadog index/retention spend, custom metrics, APM volume, and log storage
- Operate within FedRAMP High environment following change control and continuous monitoring processes
- Improve on-call operations: rotation design, escalation paths, alert quality, runbook coverage, and regional handoffs
- Partner across Support, Security, Product, and Development to ensure new services ship with monitoring, runbooks, and SLOs
Requirements:
- 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product
- Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and security fundamentals
- Production experience with an observability platform such as Datadog: metrics, logs, APM, dashboards, and monitor design; comfortable writing queries
- Demonstrated ownership of SLIs, SLOs, and error budgets for services supported
- Incident response experience: run bridges, make decisions under pressure, write postmortems
- Kubernetes in production, including ingress, deployments, resource limits, and troubleshooting
- Infrastructure as code with Terraform, plus CI/CD pipeline creation and troubleshooting (Azure DevOps preferred)
- Scripting in PowerShell and Python; fluency with YAML and JSON
- Strong networking and web fundamentals: DNS, TLS, certificate chains, load balancing, reverse proxies, firewalls, packet-level troubleshooting
- Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments
- Clear written communication for customer and executive audiences
- Willingness to participate in on-call rotation covering weekends and emergencies
- Up to 10% travel
- Must be a US citizen or have existing US work authorization; no visa sponsorship available
Nice-to-have:
- Exposure to regulated/audited environments: FedRAMP, NIST 800-53, SOC 2, ISO 27001, PCI, or similar
- AWS experience alongside Azure, including CloudFormation and SES
- Jenkins pipeline creation, SaltStack configuration management, Consul for service discovery
- Advanced log analysis in ELK stack or CloudWatch Logs Insights QL
- Web application firewall administration and bot/rate-limiting rule tuning
- Track record of controlling observability tooling spend at scale
- Microsoft Entra ID and SAML/OIDC authentication troubleshooting
- Large-scale, multi-region, geo-redundant architectures and disaster recovery testing
- Jira Service Management, PagerDuty, or similar on-call tooling experience
- Game days, chaos exercises, and mentoring in SRE practices