SlipstreamJobsFresh Startup & VC-Backed Jobs

Principal Software Engineer

Lambda - Bellevue, WA, United States - Hybrid - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers, from AI researchers to enterprises and hyperscalers. The company's mission is to make compute as ubiquitous as electricity and give everyone access to superintelligence. You will join a team of highly capable software, hardware, and network engineers building one of the largest AI training and inference networks in the world. This role reports to the VP of Cloud and AI Networking and focuses on owning the multi-year technical strategy for Lambda's networking software roadmap. Key responsibilities include: - Own the technical strategy for networking software covering backend GPU fabric, frontend networks, backbone, edge, and related infrastructure - Set technical direction and standards for other engineers to design, build, and operate against - Develop and maintain high-level technical relationships with enterprise customers on mission-critical projects - Influence across teams with strong technical judgment and resolve technical ambiguity spanning organizational boundaries - Go deep into code, network architecture, and customer use cases when needed—this is a hands-on role - Drive technical work end-to-end: shape vision, define high- and low-level design, and work alongside delivery teams - Mentor engineers, conduct design reviews, and participate in hiring to raise the bar across the organization - Participate in day-2 operations and on-call rotation to maintain the fastest feedback loop between design and customers - Serve as the senior technical voice with largest customers and translate learnings back into Lambda's roadmap - Identify new opportunities to delight customers through improved availability, performance, or new products The role requires presence in the Bellevue or San Francisco office 4 days per week, with Tuesday designated as the work-from-home day. Requirements: - 8+ years in relevant domains such as cloud computing, systems engineering, and large-scale network infrastructure - Experience designing, building, and operating production large-scale data center or cloud networks - Experience building distributed network systems and services at scale, including design, build, and operation of Tier 1 software-defined networking services - Demonstrated leadership through implementation and launch of large production-scale software services projects - Ability to functionally decompose complex problems into simple, straightforward solutions - Complete understanding of system inter-dependencies and limitations - Expert knowledge in performance, scalability, enterprise system architecture, and engineering best practices - Experience with cloud provider networking (AWS, GCP, OCI) - Excellent written and verbal communication skills; ability to leverage high-quality written artifacts to inform decision-making Nice to have: - Knowledge of HPC/AI networking: RoCEv2 and/or InfiniBand (congestion control, virtual lanes, partitions), GPUDirect RDMA, and collective communication behavior at scale - Experience with network characteristics of large distributed training workloads - Software-defined networking experience - Familiarity with network monitoring and telemetry - General awareness and ability to go deep in datacenter, backbone, and internet protocols

Similar roles