SlipstreamJobsFresh Startup & VC-Backed Jobs

ML Infrastructure Engineer

Mach9 Robotics - San Francisco, CA, United States - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mach9 Robotics is seeking an ML Infrastructure Engineer to build and maintain the systems powering production AI models for civil engineering and surveying. The company operates an ML pipeline spanning 10,000+ miles of labeled survey data, image segmentation networks, and 3D prediction models serving real-time inference to surveyors and engineers in the field. In this role, you will design and build a centralized system for versioning training data, generated datasets, and model artifacts with full lineage tracking from raw source data through trained model outputs. You'll develop and maintain reliable, reproducible ML training and data generation pipelines, refactoring existing scripts into composable, testable, and maintainable components. Key responsibilities include creating CI/CD workflows for validating data pipelines and model training runs with automated correctness checks and regression detection. You'll build tooling that enables ML engineers to launch, monitor, and debug training jobs with minimal friction. You'll also optimize and scale real-time model inference services to meet latency and throughput requirements in production, including profiling, batching strategies, and resource-efficient serving. Finally, you'll own the deployment path from trained model artifact to production endpoint, ensuring reliable rollouts, rollback, and monitoring. This role is ideal for mid-career ML infrastructure engineers with experience building for both training and inference, working with deep transformer models on hundreds of terabytes of 3D point cloud and image data, and architecting inference infrastructure that delivers both heavy offline detection algorithms and real-time responsive inference integrated with CAD software. REQUIREMENTS: - 3+ years of work experience in relevant fields - Bachelor's or Master's degree in Computer Science, Engineering, or equivalent experience - Strong communication skills and ability to work closely with ML researchers and engineers - Experience designing and building data versioning, artifact management, or dataset lineage systems (e.g., DVC, LakeFS, Weights & Biases, or custom solutions) - Hands-on experience with ML pipeline orchestration tools (e.g., Airflow, Prefect, Metaflow, or similar) - Experience with model serving and inference optimization — profiling latency, reducing memory footprint, or scaling serving infrastructure to meet real-time constraints - Ability to read and refactor ML training code; understanding of what training pipelines do well enough to make them reliable - Proficient with Python and PyTorch BONUS QUALIFICATIONS: - Familiarity with AWS infrastructure services - Experience with containerized ML workflows and GPU-accelerated training environments - Experience with model optimization techniques (e.g., quantization, TensorRT, ONNX Runtime, distillation) - Knowledge of infrastructure-as-code tools (e.g., AWS CDK, Terraform) - Experience building or operating ML systems that handle large unstructured datasets (imagery, 3D data, sensor data)

Similar roles