SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, Cloud Infrastructure

Fireworks - San Mateo, CA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fireworks is a Series D AI platform company (valued at $17.5B) that enables enterprises to build, train, and serve AI models tailored to their data and workflows. Backed by AMD, NVIDIA, Sequoia, and other top-tier investors, the company is building a virtual cloud infrastructure to serve AI workloads across multiple cloud providers globally. As a Member of Technical Staff on the Cloud Infrastructure team, you will architect and build foundational systems powering Fireworks' generative AI platform. This is a highly technical, individual-contributor role focused on distributed systems and ML infrastructure at scale. Key responsibilities include: - Architecting scalable, resilient backend infrastructure for distributed training, inference, and data processing pipelines - Designing and implementing core backend services (job schedulers, resource managers, autoscalers, model serving layers) with emphasis on efficiency and low latency - Leading technical design discussions and mentoring other engineers on best practices for large-scale ML infrastructure - Driving infrastructure optimization initiatives: compute cost reduction, storage lifecycle management, network performance tuning - Collaborating cross-functionally with ML, DevOps, and product teams to translate research needs into robust infrastructure solutions - Evaluating and integrating cloud-native technologies (Kubernetes, Kubeflow, MLFlow) to enhance platform capabilities - Owning end-to-end systems from design through deployment and observability, with strong emphasis on reliability and fault tolerance - Developing comprehensive monitoring, alerting, logging, and tracing solutions for system health and performance insights Required qualifications: 5+ years designing and building backend infrastructure in cloud environments (AWS, GCP, Azure); proven ML infrastructure experience (PyTorch, TensorFlow, Kubernetes, SageMaker); strong software development in Python or C++; deep understanding of distributed systems (scheduling, orchestration, storage, networking, compute optimization); Bachelor's in Computer Science or equivalent. Preferred: Master's/PhD; experience leading infrastructure projects for large-scale ML/AI workloads; infrastructure-as-code and CI/CD tooling (Terraform, ArgoCD); track record of performance and cost-efficiency improvements; open-source contributions.

Similar roles