SlipstreamJobsFresh Startup & VC-Backed Jobs

Engineering Manager, Fleet Engineering

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-08-22

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is hiring Engineering Managers for its Fleet Engineering organization, which owns the full lifecycle of production GPU infrastructure systems. This is a multi-team hiring initiative covering Fleet Reliability, HPC Deployments, Fleet Foundation, and Fleet Orchestration roles; candidates apply once and are matched to the best-fit team based on strengths. Fleet Engineering manages the deployment, operation, and reliability of Lambda's GPU fleet serving tens of thousands of customers. The organization spans five core functions: HPC Deployments (bare metal to production-ready capacity, firmware, burn-in, validation, InfiniBand fabric); Fleet Reliability (day-2 operations, system health, fleet uptime); Fleet Orchestration/Data (production source-of-truth systems, data synchronization, quality); Fleet Orchestration/Automation (workflow orchestration for fleet work, locking, firmware, OS installs, reporting); and Fleet Foundation (host enablement, OS provisioning, firmware management, out-of-band access, power management). In this role, you will lead and grow a distributed team of engineers responsible for deploying and operating production infrastructure. Key responsibilities include cross-functional project delivery with stakeholder alignment, identifying efficiency opportunities in tools and processes, providing visibility into progress and risks, qualifying new production technologies, managing staff allocation and priorities, conducting 1:1s and supporting career development, and participating in incident management and review programs. Required qualifications: 3+ years leading or managing engineers in AI/ML infrastructure or large-scale compute environments; ownership of production systems with real SLAs; confidence debugging across OS, hardware, and networking layers; ability to lead technical design on medium-to-large efforts; strong project management and deadline execution; effective collaboration with peer managers; deliberate team-building through hiring and performance management; excellent problem-solving instincts; and enthusiasm for hardware-software-datacenter intersection work. Nice-to-have skills include Linux systems administration, TCP/IP networking, automation and scripting, bare metal provisioning (PXE, Redfish, IPMI, BMC), strong coding ability, GPU acceleration and virtualization knowledge, datacenter infrastructure familiarity, network source-of-truth tooling (NetBox), Linux distribution building, AI-assisted development tools, customer empathy, and a technical degree or equivalent. The role requires presence in San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as work-from-home day. The work carries executive visibility and direct customer impact.

Similar roles