SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Havoc AI, founded in 2024 and headquartered in Providence, Rhode Island, is a leader in all-domain collaborative autonomy. The company develops software-defined hardware that powers military and commercial-grade autonomous systems across sea, air, and land, enabling assets to sense, decide, and act together in complex and contested environments.
As an ML Cloud Infrastructure Engineer, you will design and operate the infrastructure that enables Havoc's teams to train, evaluate, deploy, and monitor machine learning models safely and reliably. You will develop pipelines, services, platforms, and integrations connecting data lakes, telemetry stores, simulation environments, training workloads, cloud compute, and deployed models.
Key responsibilities include:
**ML Infrastructure & Pipelines:** Build pipelines transforming raw multi-modal data (telemetry, imagery, video, sensor, simulation) into curated, versioned training datasets. Develop reproducible training and evaluation workflows scaling across cloud compute and GPU resources. Build and maintain model deployment infrastructure for packaging, serving, inference, versioning, and rollback. Implement experiment tracking, dataset lineage, and model versioning. Own data schema versioning and migration as datasets and models evolve.
**Cloud Platform & Infrastructure:** Design, build, and operate scalable AWS infrastructure using Infrastructure as Code. Build and maintain Kubernetes/EKS workloads and containerized environments for training, batch processing, evaluation, and model serving. Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment. Improve utilization, scalability, and cost efficiency. Build infrastructure enabling teams to launch workloads safely.
**Reliability, Evaluation & Observability:** Build evaluation frameworks and regression testing for model quality, dataset integrity, and pipeline correctness. Establish quality and reliability signals for production readiness. Develop monitoring, logging, tracing, and observability across training jobs, data pipelines, and deployed models. Diagnose and resolve performance, scaling, and reliability bottlenecks. Maintain high standards for automation, testing, documentation, and operational readiness.
**Cross-Functional Engineering:** Partner with Autonomy, Software, Data, Simulation, and Security teams. Contribute to CI/CD and release processes. Translate engineering requirements into scalable platform capabilities. Incorporate feedback to improve ML development workflows.
**Security & Data Management:** Implement secure infrastructure practices including IAM least privilege, secrets management, and access controls. Build data and ML workflows with reproducibility, traceability, and appropriate controls. Partner with security teams to ensure ML systems meet operational and compliance requirements.
Within your first 12 months, you will have built reliable and reproducible pipelines, enabled training and evaluation workloads to run at scale with clear quality signals, delivered self-service ML infrastructure, improved observability and reliability, and established scalable foundations for growth.
**Requirements:**
- 3+ years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or related field
- Strong programming experience in Python; experience in Go, C++, or another systems-oriented language preferred
- Experience building and operating production services, APIs, data pipelines, developer platforms, or infrastructure
- Hands-on experience with ML workflows such as dataset preparation, model training, evaluation, or deployment
- Experience with cloud infrastructure, preferably AWS, and Infrastructure as Code
- Hands-on experience with Kubernetes and containerized environments
- Strong understanding of production engineering fundamentals including reliability, observability, testing, automation, and maintainability
- Ability to work effectively across engineering disciplines and solve ambiguous technical problems with high ownership
- U.S. Citizenship and ability to obtain and maintain a U.S. Government security clearance
**Nice to Have:**
- Experience with MLOps and workflow platforms such as MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster
- Experience with GPU/accelerator scheduling, distributed training, or large-scale ML workloads
- Experience with multi-modal datasets including imagery, video, telemetry, sensor, or simulation data
- Experience supporting autonomy, robotics, simulation, or real-time systems
- Experience deploying ML models to edge or embedded environments
- Experience with AWS GovCloud, GCP Assured Workloads, or compliance-driven environments such as FedRAMP or IL4/IL5