SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 212,000 - 286,000 / annual
Temporal is an open-source programming model that simplifies code, improves application reliability, and helps developers deliver features faster. The Cloud Data Store (CDS) team owns the storage, retrieval, and lifecycle management of all workflow data at planet scale.
As a Staff Software Engineer, you will design, build, and maintain significant portions of the backend for highly scalable, multi-tenant services. You'll own the custom persistence stack for Temporal Cloud, including a Write Ahead Log, metadata stores (Cassandra, etcd), multi-level caches, and tiered storage systems.
Key responsibilities include:
- Designing and building distributed data systems: craft APIs, schemas, and replication paths that keep petabytes of workflow history durable and queryable. Document design choices and operational knowledge for successful deployment and operation.
- Driving reliability and performance: own SLOs, create chaos-test plans, profile hot paths, and lead incident reviews.
- Providing technical leadership: break down roadmap epics, mentor mid-level engineers, and steward design documents through RFC processes.
- Cross-team collaboration: partner with Server, Cloud, and Developer Experience teams to deliver features end-to-end.
Required qualifications:
- 5+ years of experience building or enhancing highly scalable distributed systems
- Strong computer science fundamentals in distributed systems, multi-threading, and concurrency
- Production experience writing concurrent code in Go, Java, or similar languages at intermediate-advanced level
- Experience building and running services on AWS (Azure and GCP experience is a bonus)
- Experience with Elasticsearch and/or ClickHouse
Nice-to-have qualifications include prior contributions to Temporal, Cadence, or other workflow engines; deep expertise in storage domains (LSM trees, columnar stores, transactional logs); experience operating multi-region services with ≥99.99% uptime; open-source systems experience; and Kubernetes controller/CRD development experience.