SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 216,000 - 270,000 / annual
Scale is building reliable AI systems for critical enterprise and government decisions. The Orchestration Platform is a core internal technology that powers durable, reliable, and scalable workflows across the company—from data and model-evaluation pipelines to operational systems and customer-facing workflows.
As a Senior Software Engineer on this team, you will design and evolve the foundational orchestration platform that enables teams across Scale to author, operate, observe, and safely scale distributed workflows. This is a high-leverage infrastructure role for an engineer excited by distributed systems, reliability, and platform engineering.
Key responsibilities:
- Lead the architecture, design, implementation, and operation of Scale's core orchestration platform
- Build durable workflow infrastructure using technologies such as Temporal, Cadence, Kubernetes, and cloud-native systems
- Define platform primitives and APIs for scheduling, retries, state management, task execution, observability, and workflow lifecycle management
- Partner with product, infrastructure, data, and application teams to understand workflow needs and translate them into reusable platform capabilities
- Improve the reliability, scalability, security, and developer experience of services running critical company workflows
- Establish technical standards and best practices for distributed workflow development, deployment, testing, and incident response
- Drive cross-functional technical decisions and communicate platform direction clearly to engineers and stakeholders
Requirements:
- 5+ years of full-time software engineering experience, with focus on backend, infrastructure, and distributed systems
- Experience building and operating production systems with strong requirements for reliability, availability, and scale
- Deep familiarity with workflow orchestration platforms such as Temporal, Cadence, AWS Step Functions, Kubernetes, or similar systems
- Deep familiarity with Kubernetes and containerized production environments; familiarity with Terraform and Docker
- Strong knowledge of distributed-systems concepts including asynchronous execution, retries, idempotency, fault tolerance, state management, and observability
- Track record of leading technically complex projects from design through rollout and ongoing operation
- Excellent communication skills and ability to collaborate effectively with platform consumers and non-technical stakeholders
Nice to haves:
- Experience with data warehouses (Snowflake, Firebolt) and data pipeline/ETL tools (Dagster, dbt)
- Experience with authentication/authorization systems (Zanzibar, Authz, etc.)
- Experience scaling products at hyper-growth startups
- Excitement to work with AI technologies