SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

OneStream - Remote - Remote - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OneStream is seeking a Site Reliability Engineer to ensure the platform and services customers rely on are reliable, performant, and highly available. This role sits within the Cloud Services team and focuses on designing, implementing, and monitoring scalable and secure cloud services that power OneStream's finance operating system. You will implement application and infrastructure observability solutions, participate in on-call rotations, and partner with Product and Engineering teams to identify and maintain reliable systems. Key responsibilities include automating processes to improve reliability and performance, creating new designs and architectures for large-scale systems, sustaining high availability for key services, and mentoring team members in technical areas. You will also implement codified automated solutions integrating Dynatrace, Azure DevOps, and Jira, maintain technical documentation, and provide peer code reviews. The role requires working well in a small team, sharing responsibilities, and interacting with internal staff, managers, and customers. You will be expected to understand SOC/FedRAMP controls to assist Compliance and Security teams. Travel is not expected to exceed 5%. REQUIREMENTS: - BS/BA in computer science, engineering, or technology-related field (or equivalent work experience) - Proven work experience as a Site Reliability Engineer or similar role - 6+ years of cloud infrastructure and software development experience - 2+ years hands-on experience with Azure Kubernetes Services (AKS) with container-based deployment skills, or other platforms such as OpenShift, GKS, EKS - Advanced understanding of APM and observability tools (Dynatrace, AppInsights, DataDog, Log Analytics, New Relic, Prometheus, Grafana) - Advanced understanding of Infrastructure-as-Code (IaC) concepts and tooling (Terraform, CloudFormation templates, Bicep, ARM templates) on Microsoft Azure, AWS, or GCP - Deep knowledge of Configuration Management/Orchestration utilities (Ansible, PowerShell DSC, Chef, Puppet) - Advanced understanding of cloud concepts including elasticity, security, and identity management - Familiarity with Agile Development methodologies using Jira or Azure DevOps Boards - 6+ years hands-on experience with: automating processes using PowerShell, Bash, CLI, REST APIs, Python, ARM Templates or other scripting languages; source control tools (Git, Azure DevOps, GitHub); container orchestration platforms (Kubernetes, OpenShift, AKS, GKS, Helm); and Microsoft Azure, AWS, or Google Cloud PREFERRED: - Experience working for a cloud service provider (CSP), managed service provider (MSP), or SaaS provider - 6+ years of relevant Azure experience deploying and managing infrastructure using IaC concepts - Experience with Microsoft and .NET (.NET, C#, SQL) - Experience writing efficient and reliable code in a development environment - Linux operating system experience (Debian, Ubuntu, Alpine) - Deep knowledge of containerized applications with attention to reliability and monitoring

Similar roles