SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lightning AI, the company behind PyTorch Lightning, is seeking a Research Engineer to optimize training and inference workloads on its end-to-end AI platform. Founded in 2019 and recently merged with Voltage Park, Lightning AI combines developer-first software with cost-efficient, large-scale compute infrastructure.
You'll work at the intersection of ML systems, AI infrastructure, and performance engineering. The role is highly cross-functional, requiring deep technical problem-solving across the full stack—from model behavior and inference systems to distributed infrastructure and developer tooling.
Key responsibilities include:
- Optimizing large-scale training and inference workloads across GPUs, accelerators, and distributed systems
- Working directly with customers to analyze workloads, identify bottlenecks, and improve performance and reliability of deployed AI systems
- Developing and improving inference pipelines, model serving systems, and performance-oriented tooling for production workloads
- Designing profiling, debugging, and observability tools to analyze model execution and guide optimization strategies
- Ensuring performance improvements are accessible through clean APIs and seamless integration with the Lightning ecosystem
- Partnering with hardware vendors (NVIDIA, TPU, emerging accelerators) to support efficient execution across diverse compute backends
- Contributing to open-source projects through features, tooling, documentation, and community engagement
- Staying current with advancements in large-scale inference, distributed training, and ML systems optimization
The role is based in one of Lightning AI's hubs (NYC, SF, Seattle, or London) or remote, with a minimum of 2 in-office days per week and occasional team offsites. The company is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute, and operates globally.