SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer, Training & Experimentation

Lightning AI - San Francisco, CA, United States - Hybrid - posted 2026-08-07

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lightning AI, founded in 2019, builds an end-to-end platform for developing, training, and deploying AI systems. The company combines developer-first software with cost-efficient, large-scale compute infrastructure through its merger with Voltage Park. Lightning AI serves solo researchers, startups, and large enterprises globally, with offices in New York City, San Francisco, Seattle, and London, backed by top-tier investors including Coatue, Index Ventures, and Bain Capital Ventures. The Training & Experimentation team builds the platform infrastructure that enables developers to experiment at scale and train their own AI models. This includes support for foundation model training, fine-tuning open-source models, launching distributed training jobs, and iterating on experiments. As a Senior Software Engineer on this team, you will design and build backend systems that power AI agent orchestration and execution. Key responsibilities include developing scalable APIs and platform capabilities for tool use, workflow orchestration, memory, and state management. You'll build reliable infrastructure that enables long-running, distributed AI workflows while partnering closely with product, research, and platform engineering teams to bring new agent capabilities into production. You will improve platform reliability, observability, and performance as customer workloads scale, evaluate and strengthen technical architecture and engineering processes, and maintain high standards for software quality through testing, automation, and continuous delivery. Additionally, you will mentor engineers on system design, distributed systems, and backend engineering best practices. This role requires working across distributed training infrastructure, workload orchestration, experiment management, developer tooling, and platform APIs. The position is based in one of the San Francisco, NYC, or London office hubs with a minimum of 2 in-office days per week and occasional team and company offsites. Visa sponsorship is not available for this position.

Similar roles