SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff - Lead, Machines

Modal Labs - New York, NY, United States - In-office - posted 2026-09-22

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Modal is building a new infrastructure layer for AI—a serverless cloud platform for AI, data, and compute-intensive applications. The company recently raised $355M in Series C funding at a $4.65B valuation and has crossed $300M+ ARR, serving category-defining customers including Lovable, Ramp, Cognition, DoorDash, and Suno. You will lead the team responsible for Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. This is a hands-on technical leadership role where you'll manage a team of 3–8 engineers while staying deeply involved across the full stack. Your responsibilities include: - Owning the full lifecycle of machines, from accepting and benchmarking new hardware from multiple providers to network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation of unhealthy hosts - Managing a team of engineers through project planning, growth, and performance conversations - Staying hands-on across the stack: BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services - Shaping the long-term technical direction and driving architectural decisions - Participating in on-call rotation and responding to production incidents Key initiatives the team is working on include automatic remediation of unhealthy machines, automatic integration of new hardware into the fleet, network health monitoring across datacenters, automatic hardware acceptance testing and benchmarking, and custom network bootloader and machine image pipeline management. REQUIREMENTS: - 7+ years of experience writing high-quality production code - 3+ years of direct people management experience, ideally leading a team of engineers through project planning, growth, and performance conversations - Experience operating large fleets of physical hardware at scale (bare metal provisioning, BMC/IPMI, PXE and network boot, firmware) or building the control planes that manage them - Strong cloud skills - Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers, etc.) - Experience working with hardware and colocation providers, including hardware acceptance testing and benchmarking - Track record of setting technical direction and driving architectural decisions across a team - Willingness to step into on-call rotation and respond to production incidents NICE-TO-HAVES: - Experience with GPUs and the NVIDIA software stack (drivers, health monitoring, RDMA/NVLink) in production - Prior experience with Go

About Modal Labs

AI / Data / Infrastructure — serverless cloud platform for AI, data, and compute-intensive applications.

Similar roles