SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Camunda - Remote - Remote - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 149,800 - 241,500 / annual

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end business processes. The company is trusted by over 700 organizations worldwide, including 9 of the top 10 US banks, and is named on GP Bullhound's Top 100 Next Unicorn list and recognized as a Visionary in the 2025 Gartner Magic Quadrant for Business Orchestration and Automation Technologies. As a Senior Site Reliability Engineer, you will design and maintain Camunda's Kubernetes-based multi-cloud platform, ensuring it is available, scalable, and fault-tolerant. You will establish configuration best practices and network services that engineering teams rely on daily. Your responsibilities include implementing and improving monitoring and alerting tools that give both SREs and developers real visibility into system health and performance across the stack. You will own systems end-to-end, adopting a "you build it, you run it" mentality, which includes participating in on-call rotations and being the go-to person for quick fixes during incidents. You'll create runbooks and automation that turn complex problems into manageable processes. Working cross-functionally with product engineering, product management, and support, you'll define and deliver infrastructure features that move the needle. You will identify repetitive work and automate it away, sharing learnings with teammates to elevate the entire organization's infrastructure capabilities. As a senior engineer, you will mentor less experienced engineers on complex infrastructure challenges, breaking down technical problems into clear steps and supporting others in growing their skills. You'll leverage AI tools responsibly for research, code review, documentation, test generation, and automation to improve your effectiveness, while maintaining human accountability for all decisions and deliverables. REQUIREMENTS: Must Haves: - Deep hands-on experience with Kubernetes: built, deployed, and maintained Kubernetes clusters in production environments; understand workload, networking, and storage management at scale - Infrastructure as code expertise: fluent in Terraform or similar IaC tools; able to version, test, and safely deploy infrastructure changes - Demonstrated experience in monitoring and observability: worked with Prometheus, Grafana, or similar platforms to instrument systems and alert on what matters - Strong 3rd-level support and incident response skills: responded to production incidents, diagnosed complex issues, communicated clearly with stakeholders under pressure; understand root cause analysis and prevention - Passion for automation and raising the quality bar: see manual work as a problem to solve; care deeply about building reliable, maintainable, easy-to-understand systems - Responsible use of AI tools: leverage AI for research, code review, documentation, test generation, and automation; validate AI outputs against requirements; never share confidential or personal data with AI systems without authorization; retain human accountability for decisions and deliverables Nice-to-Haves: - Experience with major cloud providers: AWS (EKS), Google Cloud Platform (GKE), or similar managed Kubernetes services - ArgoCD or GitOps workflows: used declarative, Git-driven infrastructure management - Proficiency in Python, Go, or similar languages: write scripts and automation tools that solve real problems - Experience with SLOs and alerting frameworks: helped teams define meaningful service level objectives and set up alerts that don't cry wolf

Similar roles