SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer, Reliability

Roblox - San Mateo, CA, United States - Hybrid - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 243,290 - 295,250 / annual

Roblox is seeking a Senior Software Engineer for its Reliability team to drive the evolution of systems supporting tens of millions of daily active users. The role focuses on building robust infrastructure that meets the highest standards of performance, reliability, and efficiency at scale. Key Responsibilities: - Design and implement libraries that promote fault-tolerance and resilience, including retries, circuit breakers, and adaptive concurrency limits - Build, automate, and standardize process automation to create a "golden path" of tooling and platform support for the Roblox ecosystem - Develop tooling that provides production guardrails, such as evaluating release candidate capacity with load testing before production deployment - Create performance monitoring services and observability tools to understand capacity issues and platform degradations - Build monitoring tooling for production services and their changes, including generalized canarying services with alerting - Collaborate with cross-functional teams to solve complex technical challenges at scale You will work on infrastructure that powers a platform aiming to reach 1 billion daily active users, with a focus on reliability, scalability, and developer experience. Workplace: Hybrid (onsite Tuesday, Wednesday, Thursday; optional Monday and Friday). Requirements: - BS degree in Computer Science or related engineering field, or equivalent professional experience - Minimum 4+ years of software engineering experience, with preference for Site Reliability Engineering (SRE) or related background - Experience writing common programming languages (Go, C#, Java, or similar) - Demonstrated experience building software and tools with broad adoption across teams - Strong systems thinking and focus on code reliability - Experience with large project lifecycles, sprint planning, and breaking down complex tasks into milestones - Ability to approach problems with curiosity, understand issues deeply before coding, and use data to validate theories - Self-organized approach to complex problems with ability to overcome emergent issues

Similar roles