SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fireworks is a Series D AI platform company (valued at $17.5B) that enables enterprises to build, train, and serve specialized AI models tailored to their data and workflows. Backed by AMD, NVIDIA, Sequoia, Benchmark, and other tier-one investors, the company powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads.
As a Training Infrastructure Engineer, you will design, develop, and maintain large-scale backend and cloud-native infrastructure supporting distributed machine learning training, inference, and data processing pipelines. This is a senior individual contributor role with mentorship and technical leadership responsibilities.
Key responsibilities include:
- Architect and build scalable, resilient backend infrastructure for distributed training, inference, and data processing
- Lead technical design discussions and mentor engineers on large-scale ML systems best practices
- Design and implement core backend services optimized for efficiency and low latency
- Drive infrastructure optimization for compute cost, storage lifecycle, and network performance
- Collaborate with ML, DevOps, and product teams to translate research and product requirements into robust solutions
- Evaluate and integrate cloud-native technologies (Kubernetes, Ray, Kubeflow, MLFlow)
- Own end-to-end systems from design to deployment, emphasizing reliability and operational excellence
Required qualifications: Bachelor's in Computer Science or equivalent plus 4+ years in software engineering. Specifically: 4+ years designing and optimizing large-scale backend infrastructure and distributed data systems (PostgreSQL, MySQL, DynamoDB, Spark, Flink, Kafka) in cloud environments (AWS, GCP, Azure); 4+ years with server-side languages (Python, C++, Go, TypeScript); 4+ years writing technical design docs and leading cross-functional projects; 3+ years developing data processing and API systems (gRPC, Thrift); 3+ years with A/B testing and experimentation platforms; 3+ years conducting coding interviews; 2+ years with Docker and Kubernetes; 2+ years defining data-driven metrics.