SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fireworks is a Series D AI platform company (valued at $17.5B) that enables enterprises to build, train, and serve specialized AI models tailored to their data and workflows. Backed by AMD, NVIDIA, Sequoia, Benchmark, and other tier-one investors, the company powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads.
As a Member of Technical Staff in AI Training Infrastructure, you will design, build, and optimize the infrastructure powering large-scale model training operations. This is a senior individual contributor role focused on solving hard problems at the forefront of AI infrastructure.
Key responsibilities include:
- Design and implement scalable infrastructure for large-scale LLM and multimodal model training workloads
- Develop and maintain distributed training pipelines, optimizing performance across multiple GPUs, nodes, and data centers
- Implement monitoring, logging, and debugging tools for training operations
- Architect and maintain data storage solutions for massive training datasets
- Automate infrastructure provisioning, scaling, and orchestration using Kubernetes and related tools
- Collaborate with AI researchers to implement and optimize training methodologies
- Analyze and improve efficiency, scalability, and cost-effectiveness of training systems
- Troubleshoot complex performance issues in distributed training environments
Minimum qualifications: Bachelor's in Computer Science or equivalent; 3+ years with distributed systems and ML infrastructure; PyTorch experience; proficiency in AWS/GCP/Azure; containerization and Kubernetes expertise; knowledge of distributed training techniques (data parallelism, model parallelism, FSDP).
Preferred: Master's/PhD; experience training large language models; ML workflow orchestration tools; high-performance distributed computing optimization; ML DevOps practices; open-source ML infrastructure contributions.
This role offers ownership and impact in a fast-growing team building bleeding-edge AI infrastructure technology with minimal bureaucracy.