SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Infrastructure Software Engineer

Lightning AI - New York, NY, United States - Hybrid - posted 2026-08-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lightning AI is seeking a Senior Infrastructure Software Engineer to join its Infrastructure Engineering team. The company, founded in 2019 and behind PyTorch Lightning, builds an end-to-end platform for developing, training, and deploying AI systems. Through a merger with Voltage Park, Lightning AI combines developer-first software with cost-efficient, large-scale compute infrastructure. The Infrastructure Engineering team builds the software systems that operate and manage Lightning AI's large-scale GPU and bare-metal infrastructure. You will develop services, APIs, tooling, and automation that bridge the gap between physical infrastructure and the systems customers depend on for AI/ML training, inference, and HPC workloads. In this role, you will build production software and automation supporting infrastructure throughout its lifecycle—from provisioning and bring-up through monitoring, maintenance, and decommissioning. You will design systems enabling engineers and software to interact programmatically with thousands of servers, storage systems, and high-performance networks, reducing manual operations and making infrastructure management more reliable, repeatable, and scalable. Key responsibilities include designing and building production services, APIs, tooling, and automation for managing large-scale bare-metal and GPU infrastructure; building and extending software systems supporting provisioning, configuration, monitoring, and lifecycle management; developing reliable automation that reduces manual work and improves consistency and scalability; and building tools enabling programmatic interaction with physical infrastructure. This is a software-first infrastructure role combining strong software engineering with Linux systems and hardware-aware infrastructure automation. You will work closely with teams across Infrastructure, Networking, Data Center Operations, and Platform Engineering to bring new capacity online and continuously improve infrastructure operations at scale. The role is based in one of Lightning AI's hubs (NYC, SF, Seattle, or London) with a minimum of 2 in-office days per week and occasional team and company offsites.

Similar roles