SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Scopely is hiring a Senior Platform Engineer (Reliability) to join an unannounced multiplayer strategy/MMO game in early development. You'll own the reliability and operational excellence of distributed backend systems, with a primary focus on messaging infrastructure (NATS/NATS JetStream).
Key responsibilities include:
**Messaging Systems Ownership**: Operate, monitor, and improve NATS cluster and JetStream infrastructure, ensuring reliability, observability, and team understanding of the system.
**Observability & Signals**: Design the observability layer for distributed backend systems—defining metrics, traces, and logs that make operational problems visible and actionable across teams.
**Cross-functional Diagnosis**: Correlate infrastructure signals (IOPS, latency, resource saturation) with application behavior (message throughput, consumer lag, retry storms) to identify root causes and guide teams toward solutions.
**SLOs & Error Budgets**: Define, implement, and maintain SLO frameworks for backend services and messaging pipelines, helping teams balance velocity and stability.
**Reliability as Internal Product**: Build tooling, runbooks, and operational frameworks that enable backend and infrastructure engineers to self-serve on operational concerns, reducing dependency on SRE.
**Incident Management**: Lead or contribute to incident response, drive postmortems toward systemic fixes, and embed preventative improvements into engineering workflows.
**Code Literacy & Collaboration**: Navigate infrastructure and backend codebases confidently; identify instrumentation gaps, architectural risks, and operational anti-patterns in close collaboration with engineering teams.
Required experience: Strong SRE, production operations, or backend engineering background with operational focus. Hands-on production experience with NATS/NATS JetStream or equivalent distributed messaging systems (Kafka, Kinesis). Ability to read and reason about application code (C#, Go, Python, etc.). Strong observability design and implementation skills. Experience debugging complex distributed systems across infrastructure and application boundaries. Solid cloud infrastructure knowledge (AWS preferred) and containerized workloads (ECS, Kubernetes/EKS). Strong communication skills translating between infrastructure and application engineers. Infrastructure as Code first mindset.
Bonus: Direct NATS JetStream experience, Datadog or equivalent observability platform, database operational concerns, mentoring experience.