SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Modal is building a new infrastructure layer for AI—a serverless cloud platform for AI, data, and compute-intensive applications. The company recently raised $355M in Series C funding at a $4.65B valuation and has crossed $300M+ ARR. Customers include Lovable, Ramp, Cognition, DoorDash, and Suno.
You will work on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, plus the control plane that provisions, images, monitors, and repairs them. Your responsibilities include:
- Designing and building high-performance systems for the serverless platform
- Automating the integration of new capacity from a growing set of hardware providers
- Auditing and benchmarking hosts and clusters
- Maintaining machine images, configuring GPUs, RDMA, networking, and storage
- Getting machines into production at scale
- Building automation to keep the fleet healthy without human intervention
- Detecting and remediating bad GPUs, thermal issues, and disk failures
- Debugging across layers: kernel panics, broadcast storms during boot, container runtime compatibility on new architectures
- Participating in on-call rotation and responding to production incidents
Key team initiatives include automatic remediation of unhealthy machines, automatic integration of new CPU/GPU/storage servers while managing hardware heterogeneity, network health monitoring across datacenters, automatic hardware acceptance testing and benchmarking, and custom network bootloader and machine image pipelines.
REQUIREMENTS:
- 5+ years of experience writing high-quality production code
- Experience operating fleets of physical hardware (bare metal provisioning, BMC/IPMI, PXE or network boot, firmware) or building control planes that manage them
- Strong cloud skills
- Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers)
- Effective debugging across layers (BGP, Linux RPS, vBIOS, Python services)
- Willingness to participate in on-call rotation and respond to production incidents
NICE-TO-HAVES:
- Experience with GPUs and NVIDIA software stack in production (drivers, health monitoring, XIDs, RDMA/NVLink)
- Prior experience with Go
About Modal Labs
AI / Data / Infrastructure — serverless cloud platform for AI, data, and compute-intensive applications.