SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, AI Training Infrastructure

Fireworks - San Mateo, CA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fireworks is a Series D AI platform company (valued at $17.5B) that enables enterprises to build, train, and serve specialized AI models tailored to their data and workflows. Backed by AMD, NVIDIA, Sequoia, Benchmark, and other tier-one investors, the company powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. As a Member of Technical Staff in AI Training Infrastructure, you will design, build, and optimize the infrastructure powering large-scale model training operations. This is a senior individual contributor role focused on solving hard problems at the forefront of AI infrastructure. Key responsibilities include: - Design and implement scalable infrastructure for large-scale LLM and multimodal model training workloads - Develop and maintain distributed training pipelines, optimizing performance across multiple GPUs, nodes, and data centers - Implement monitoring, logging, and debugging tools for training operations - Architect and maintain data storage solutions for massive training datasets - Automate infrastructure provisioning, scaling, and orchestration using Kubernetes and related tools - Collaborate with AI researchers to implement and optimize training methodologies - Analyze and improve efficiency, scalability, and cost-effectiveness of training systems - Troubleshoot complex performance issues in distributed training environments Minimum qualifications: Bachelor's in Computer Science or equivalent; 3+ years with distributed systems and ML infrastructure; PyTorch experience; proficiency in AWS/GCP/Azure; containerization and Kubernetes expertise; knowledge of distributed training techniques (data parallelism, model parallelism, FSDP). Preferred: Master's/PhD; experience training large language models; ML workflow orchestration tools; high-performance distributed computing optimization; ML DevOps practices; open-source ML infrastructure contributions. This role offers ownership and impact in a fast-growing team building bleeding-edge AI infrastructure technology with minimal bureaucracy.

Similar roles