SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
SpAItial is building next-generation world models that natively understand the physics and geometry of 3D environments, applying generative AI and computer vision to robotics, AR/VR, gaming, and cinema. The company is seeking a Machine Learning Systems & Infrastructure Engineer to design and operate the scalable systems that transform raw real-world data into trained world models and production endpoints.
You will own and evolve the ML systems enabling training, evaluation, and serving of large foundation models. Key responsibilities include:
**Distributed Training & Performance**: Improve high-throughput training stacks (PyTorch DDP/FSDP, NCCL) for performance, stability, and reproducibility. Implement preemption-safe and sharded checkpointing strategies.
**Data Systems & Pipelines**: Build end-to-end Python pipelines ingesting third-party capture sources into clean, versioned training datasets. Handle scraping (Playwright), preprocessing, and optimize petabyte-scale storage using object storage, FUSE mounts, caching layers, and metadata stores.
**ML Workflow Orchestration & Serving**: Operate systems researchers use to launch experiments and production endpoints. Work with workflow engines (Kubeflow Pipelines, Airflow), GPU schedulers (Volcano, Slurm), experiment trackers (MLflow, Weights & Biases), and inference platforms (Modal, Triton). Maintain a launcher SDK for streamlined runs.
**Infrastructure & Reliability**: Ship workloads via Docker and Kubernetes. Maintain IaC (Terraform) and CI/CD pipelines including self-hosted GPU runners. Implement monitoring, logging, and alerting (Prometheus/Grafana, OpenTelemetry) with defined SLOs and incident response.
**Security & Collaboration**: Manage secrets, IAM, and network boundaries. Partner with ML researchers and engineers to unblock work and improve developer experience.
You'll work hands-on and code-heavy in a shared monorepo with the research team, primarily in Python. The role demands strong software engineering fundamentals, 3+ years writing production Python in large codebases, and deep experience with modern ML training stacks, distributed debugging, and large-scale data pipelines. GPU compute expertise, cloud proficiency (AWS/GCP/Azure), containerization, IaC, and observability tooling are essential.