SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers, from AI researchers to enterprises and hyperscalers. The company's mission is to make compute as ubiquitous as electricity and give everyone access to superintelligence.
As a Senior Software Engineer in Lambda's Cloud Services Engineering organization, you will design, build, and operate the distributed systems powering Lambda's GPU cloud. Your team owns platform capabilities across compute control planes, managed Kubernetes, cloud APIs, identity and access, usage metering and billing, capacity and orchestration, reliability, and developer-facing infrastructure.
You will be a full-cycle engineer, owning systems through design, deployment, on-call, incident follow-through, and continuous improvement. You will turn large-scale GPU infrastructure into reliable, secure, customer-facing cloud services by building APIs, workflows, stateful controllers, schedulers, and operational tooling.
Key responsibilities include:
- Design, build, and operate services, APIs, control planes, and platform capabilities that power Lambda's AI cloud
- Solve distributed-systems problems involving state, consistency, concurrency, scheduling, failure recovery, and safe lifecycle management
- Own the full engineering lifecycle: problem framing, architecture, implementation, testing, rollout, observability, on-call, and continuous improvement
- Improve system availability, latency, throughput, efficiency, security, and operability as Lambda scales
- Turn incidents and near misses into durable engineering improvements through better automation, testing, guardrails, and backstops
- Work across product, infrastructure, networking, storage, security, and SRE teams to resolve dependencies
- Use AI-assisted development tools with judgment to accelerate exploration while independently verifying correctness, security, and maintainability
- Contribute to technical standards, design and code reviews, and mentorship
This role is a strong fit for engineers who enjoy cloud infrastructure, distributed systems, operational excellence, and solving ambiguous problems across software and infrastructure boundaries.
Requirements:
- 7+ years of professional software engineering experience, or equivalent evidence of impact building production systems
- Depth in at least one general-purpose language (Go and Python are primary); ability to reason about concurrency, error handling, and testing
- Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at meaningful scale
- Practical understanding of system design, data models, APIs, failure modes, performance, and tradeoffs required to run reliable software in production
- Track record of owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement
- Proven ability to align cross-functional partners and gain consensus around decisions and tradeoffs
Nice to have:
- 2+ years building cloud services or platform infrastructure, or operating large-scale production systems on AWS, GCP, Azure, or comparable cloud platform
- Experience with Kubernetes, container orchestration, schedulers, controllers, or cloud control-plane systems
- Depth in cloud infrastructure or platform domains such as compute, storage, networking, identity and access, developer platforms, container orchestration, usage metering and billing, databases, or fleet management
- Experience with infrastructure automation, durable workflow systems, event-driven architectures, or infrastructure as code
- Experience designing highly available, multi-region, or rapidly scaling distributed systems
- Familiarity with GPU infrastructure, HPC environments, or large-scale AI/ML training and inference workloads