SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fireworks is a Series D AI platform company (valued at $17.5B) that enables enterprises to build, train, and serve AI models tailored to their data and workflows. Backed by AMD, NVIDIA, Sequoia, and other top-tier investors, the company is building a virtual cloud infrastructure to serve AI workloads across multiple cloud providers globally.
As a Member of Technical Staff on the Cloud Infrastructure team, you will architect and build foundational systems powering Fireworks' generative AI platform. This is a highly technical, individual-contributor role focused on distributed systems and ML infrastructure at scale.
Key responsibilities include:
- Architecting scalable, resilient backend infrastructure for distributed training, inference, and data processing pipelines
- Designing and implementing core backend services (job schedulers, resource managers, autoscalers, model serving layers) with emphasis on efficiency and low latency
- Leading technical design discussions and mentoring other engineers on best practices for large-scale ML infrastructure
- Driving infrastructure optimization initiatives: compute cost reduction, storage lifecycle management, network performance tuning
- Collaborating cross-functionally with ML, DevOps, and product teams to translate research needs into robust infrastructure solutions
- Evaluating and integrating cloud-native technologies (Kubernetes, Kubeflow, MLFlow) to enhance platform capabilities
- Owning end-to-end systems from design through deployment and observability, with strong emphasis on reliability and fault tolerance
- Developing comprehensive monitoring, alerting, logging, and tracing solutions for system health and performance insights
Required qualifications: 5+ years designing and building backend infrastructure in cloud environments (AWS, GCP, Azure); proven ML infrastructure experience (PyTorch, TensorFlow, Kubernetes, SageMaker); strong software development in Python or C++; deep understanding of distributed systems (scheduling, orchestration, storage, networking, compute optimization); Bachelor's in Computer Science or equivalent.
Preferred: Master's/PhD; experience leading infrastructure projects for large-scale ML/AI workloads; infrastructure-as-code and CI/CD tooling (Terraform, ArgoCD); track record of performance and cost-efficiency improvements; open-source contributions.