SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer - FedRAMP

Delinea - Remote - Remote - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Delinea is seeking a Senior Site Reliability Engineer to join the Cloud Engineering team and own the availability and performance of production SaaS services running in Azure and AWS, including a FedRAMP High environment. This is a hands-on role focused on measuring reliability, tuning detection, responding to incidents, and automating manual work. Key responsibilities include: - Own reliability end-to-end for a set of production SaaS services: availability, performance, and capacity planning - Define and manage service-level indicators (SLIs), service-level objectives (SLOs), and error budgets to prioritize reliability work - Build and tune monitoring in Datadog and Azure Monitor, including threshold, composite, and anomaly detection monitors, synthetic checks, dashboards, and alert routing - Automate incident response by wiring monitors into remediation workflows and replacing manual runbook steps with code - Participate in on-call rotation, lead incident response for high-severity events, and coordinate resolution across teams - Write post-incident reviews and customer-facing root cause analyses, then drive preventive actions to completion - Work support escalations: reproduce issues, diagnose from logs/traces/network captures, and feed patterns back to product - Build and maintain infrastructure as code using Terraform and Azure DevOps pipelines - Administer the web application firewall: rule tuning, rate limiting, false positive triage, and security team coordination - Own observability platform costs: manage Datadog index/retention spend, custom metrics, APM volume, and log storage - Operate within FedRAMP High environment following change control and continuous monitoring processes - Improve on-call operations: rotation design, escalation paths, alert quality, runbook coverage, and regional handoffs - Partner across Support, Security, Product, and Development to ensure new services ship with monitoring, runbooks, and SLOs Requirements: - 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product - Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and security fundamentals - Production experience with an observability platform such as Datadog: metrics, logs, APM, dashboards, and monitor design; comfortable writing queries - Demonstrated ownership of SLIs, SLOs, and error budgets for services supported - Incident response experience: run bridges, make decisions under pressure, write postmortems - Kubernetes in production, including ingress, deployments, resource limits, and troubleshooting - Infrastructure as code with Terraform, plus CI/CD pipeline creation and troubleshooting (Azure DevOps preferred) - Scripting in PowerShell and Python; fluency with YAML and JSON - Strong networking and web fundamentals: DNS, TLS, certificate chains, load balancing, reverse proxies, firewalls, packet-level troubleshooting - Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments - Clear written communication for customer and executive audiences - Willingness to participate in on-call rotation covering weekends and emergencies - Up to 10% travel - Must be a US citizen or have existing US work authorization; no visa sponsorship available Nice-to-have: - Exposure to regulated/audited environments: FedRAMP, NIST 800-53, SOC 2, ISO 27001, PCI, or similar - AWS experience alongside Azure, including CloudFormation and SES - Jenkins pipeline creation, SaltStack configuration management, Consul for service discovery - Advanced log analysis in ELK stack or CloudWatch Logs Insights QL - Web application firewall administration and bot/rate-limiting rule tuning - Track record of controlling observability tooling spend at scale - Microsoft Entra ID and SAML/OIDC authentication troubleshooting - Large-scale, multi-region, geo-redundant architectures and disaster recovery testing - Jira Service Management, PagerDuty, or similar on-call tooling experience - Game days, chaos exercises, and mentoring in SRE practices

Similar roles