SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fireworks is a Series D AI infrastructure platform enabling companies to build, train, and serve specialized AI models tailored to their data and workflows. Founded by the PyTorch team and backed by major investors including AMD, NVIDIA, Sequoia, and Benchmark Capital, Fireworks is valued at $17.5 billion and operates a production AI platform serving hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads.
As a Software Engineer on the AI Infrastructure team, you will design and develop core systems powering Fireworks' generative AI platform. Your focus will be on building reliable, performant, and user-friendly infrastructure that bridges customer needs with Fireworks' proprietary inference engine.
Key responsibilities include:
- Design and develop scalable backend infrastructure supporting distributed training, inference, and data pipelines
- Build and maintain core backend services including LLM CI/CD pipelines, control plane, and model serving systems
- Optimize performance, cost efficiency, and reliability across compute, storage, and networking layers
- Develop frameworks and safeguards ensuring industry-leading model quality
- Collaborate with performance, training, and product teams to translate research and product requirements into infrastructure solutions
- Participate in code reviews, technical discussions, and continuous integration/deployment processes
Minimum qualifications: Bachelor's in Computer Science or related field (or equivalent experience); 3+ years software engineering experience focused on infrastructure or ML systems; strong programming skills in Python, Go, or similar languages; proven ML infrastructure and tooling experience (PyTorch, MLflow, Vertex AI, SageMaker, Kubernetes); basic LLM knowledge (context length, prefill, KV cache memory estimation).
Preferred: 5+ years infrastructure/ML systems experience; experience with open-source inference engines (vLLM, Sglang, TRT-LLM); open-source contributions; large-scale ML/MLOps infrastructure building experience.