SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Havoc AI is building software-defined autonomous systems for military and commercial applications across sea, air, and land. As a Data & ML Infrastructure Engineer, you will own the data infrastructure that enables the company to develop, evaluate, and continuously improve these autonomous systems.
You will design and build scalable pipelines that ingest, store, index, and organize large volumes of video, imagery, telemetry, sensor data, autonomy logs, and mission data. Your work will transform messy real-world operational data into organized, searchable, reproducible datasets that power model training, system evaluation, failure reproduction, and capability improvement.
Key responsibilities include:
**Data Infrastructure & Pipelines**: Build and maintain infrastructure for multimodal operational data. Own data ingestion, storage, indexing, metadata management, access patterns, and lifecycle management within Havoc's data lake. Develop scalable pipelines that transform raw data into curated datasets for ML training, evaluation, debugging, and analysis. Build tools for searching, filtering, tagging, and retrieving data across platforms, missions, and operating conditions.
**Dataset Curation & ML Enablement**: Create workflows to select, clean, label, validate, and version datasets. Partner with Autonomy, Perception, Software, and Field Operations teams to identify high-value data for model development. Support annotation and labeling workflows. Develop reproducible dataset-generation workflows for training, validation, regression testing, and benchmarking. Integrate datasets with model training, experiment tracking, evaluation, and deployment workflows. Support multimodal dataset construction including sensor synchronization and alignment.
**Data Quality & Reliability**: Develop automated checks for missing streams, corrupted files, synchronization issues, metadata gaps, and pipeline failures. Establish standards for dataset quality, lineage, versioning, and reproducibility. Build monitoring and observability around critical pipelines. Troubleshoot complex data and infrastructure issues. Use field data and logs to help teams understand system performance.
**Developer Tools & Collaboration**: Build self-service tools that make operational data easier for engineers to discover and use. Partner closely with cross-functional teams. Translate engineering and ML requirements into scalable data capabilities. Improve workflows for replaying, visualizing, analyzing, and comparing operational data. Maintain documentation, data standards, and best practices.
Within 12 months, success means: reliable pipelines moving operational data into organized storage; significantly easier data discovery and access for engineers; reproducible workflows for high-quality datasets; improved data quality, lineage, and observability; and faster iteration from field data to model improvement to deployment.
**Requirements:**
- Bachelor's degree in Computer Science, Data Science, Machine Learning, Electrical Engineering, Computer Engineering, Robotics, Applied Mathematics, or related technical field
- 3+ years of experience in data engineering, ML infrastructure, data platforms, backend systems, MLOps, or related engineering roles
- Experience designing and operating production data pipelines for large-scale structured, semi-structured, or unstructured datasets
- Experience working with video, imagery, time-series telemetry, sensor data, logs, or other high-volume operational data
- Strong programming skills in Python and SQL
- Experience with cloud storage, object stores, data lakes, databases, distributed processing, or modern data platforms
- Familiarity with dataset versioning, metadata management, data lineage, access controls, and reproducible data workflows
- Strong software engineering fundamentals including testing, reliability, maintainability, and observability
- Strong debugging skills and comfort working across complex data pipelines and production infrastructure
- Ability to operate independently and take ownership in a fast-moving engineering environment
- U.S. citizenship and ability to obtain and maintain a U.S. Government security clearance
**Nice to Have:**
- Experience with ML infrastructure, MLOps, training pipelines, experiment tracking, model evaluation, or model registries
- Experience managing video, perception, telemetry, or autonomous-system datasets
- Experience with S3-compatible storage, PostgreSQL, Spark, Ray, Airflow, Dagster, Kubernetes, Docker, or Kafka
- Experience with data catalogs, dataset versioning platforms, feature stores, or labeling tools
- Experience building search, replay, visualization, or analysis tools for video, telemetry, logs, or sensor data
- Experience supporting annotation workflows for computer vision, perception, tracking, or autonomy
- Familiarity with sensor synchronization, timestamp alignment, calibration metadata, log replay, or multimodal dataset construction
- Experience with security, access controls, auditability, and data-handling requirements in government or defense environments
- Experience supporting defense, robotics, autonomy, aerospace, or dual-use technology programs
- Active or prior security clearance