SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is building the Superintelligence Cloud, a leader in AI cloud infrastructure serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. The company's mission is to make compute as ubiquitous as electricity and democratize access to superintelligence.
You will join a team of highly capable software, hardware, and network engineers building one of the largest AI training and inference networks in the world. This role reports to the VP of Cloud and AI Networking and offers the opportunity to own Lambda's multi-year technical strategy for network infrastructure—fabric, backbone, and edge—while continually reinventing the current topology, routing design, and automation.
Key responsibilities include:
- Owning the multi-year technical strategy for Lambda's entire network infrastructure
- Setting technical direction and standards that other engineers design, build, and operate against
- Owning high-level technical relationships with enterprise customers on mission-critical projects
- Influencing across organizational boundaries (network, hardware, platform, product teams) to resolve technical ambiguity
- Going deep on the hardest problems—packet paths, routing tables, vendor firmware, code—with hands-on involvement
- Driving technical work end-to-end: shaping vision, defining high- and low-level design, working alongside delivery teams
- Mentoring engineers, conducting design reviews, and raising the bar through hiring
- Participating in day-2 operations and on-call rotation to maintain the fastest feedback loop between design and reality
- Serving as the senior technical voice with largest customers on mission-critical projects
- Identifying new opportunities to delight customers through novel improvements to availability, performance, or new products
The role is hybrid, requiring presence in the Bellevue or San Francisco office 4 days per week; Tuesday is Lambda's designated work-from-home day.
REQUIREMENTS:
- 8+ years in relevant domains such as cloud computing, systems engineering, and large-scale network infrastructure
- Experience designing, building, and operating production large-scale data center or cloud networks
- Experience designing, building, and operating software-defined networking services and distributed systems at the highest level of availability
- Track record leading through implementation of large production-scale networking projects
- Experience building distributed network systems and services at scale
- Experience with cloud provider networking (AWS, GCP, OCI)
- Deep expertise in datacenter, backbone, and internet protocols and technology
- Experience building and optimizing networks for best customer experience (lowest latency, fastest convergence, highest availability)
- Comfortable on Linux command line with understanding of Linux networking stack and internals
- Production automation skills in Python and a configuration management tool, working with network APIs and source control
- Production experience with multiple network gear vendors (Arista, Juniper, Cisco, Cumulus/SONiC, Opengear)
- Experience with networking monitoring stacks (Datadog, Clickhouse, Grafana, Prometheus, gNMI, OTel)
- Excellent written and verbal communication skills; ability to leverage high-quality written artifacts for decision-making
NICE TO HAVE:
- Hands-on experience with HPC/AI networking: RoCEv2 and/or InfiniBand (congestion control, virtual lanes, partitions), GPUDirect RDMA, collective communication at scale
- Experience with network characteristics of large distributed training workloads
- Experience automating network configuration in public cloud providers using Terraform, Ansible, or Salt
- Software-defined networking experience
- Modern network monitoring and telemetry (Prometheus, Grafana, gNMI, OpenTelemetry)
- Source-of-truth and IPAM tooling experience (Netbox)
- DWDM and SD-WAN experience
- Understanding of data center power, space, and cooling trade-offs and topology impact
- Virtualization technology and load balancer experience