SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 170,000 - 300,000 / annual
Hadrian is building autonomous factories to reindustrialize American manufacturing, combining AI, advanced software, robotics, and full-stack manufacturing to help aerospace and defense companies build rockets, satellites, aircraft, and mission-critical systems up to 10x faster and at lower cost. The company recently closed a $1.37B Series D at a $7.87B valuation and is rapidly expanding manufacturing capabilities across welding, casting, forging, electronics, and additive manufacturing.
As an ML Platform Engineer, you will own the production infrastructure that enables Hadrian's factories to safely depend on machine learning models for drawing extraction, cycle-time prediction, forecasting, scheduling, and more. You'll standardize deployment patterns built around MLflow, Dagster, ECR, FastAPI, and EKS, making them the backbone for packaging, evaluating, releasing, serving, monitoring, and rolling back models across the manufacturing footprint.
Key responsibilities include: building production platforms that ensure models remain reliable, effective, and secure; developing shared batch and online serving for tabular, vision, document-AI, scheduling, graph, and embedding workloads with clear SLAs; creating repeatable release and evaluation processes with automated tests, reproducible artifacts, lineage, shadow deployments, canaries, and A/B tests; owning online feature serving and maintaining contract integrity with offline feature tables; building operational tooling for telemetry, incident response, autoscaling, resource and GPU management, cost attribution, and secure model routing; and developing APIs, SDKs, reusable templates, and documentation for adoption by other teams.
Required qualifications: track record building and operating production ML infrastructure across multiple models or inference workloads; strong production-level Python and SQL skills including typing, testing, packaging, API design, and observability; hands-on experience with Kubernetes, containers, and distributed-system failure modes; engineering background with model registries, feature systems, batch/real-time inference, experiment tracking, or model CI/CD workflows; practical judgment around latency, throughput, availability, multi-tenancy, autoscaling, and infrastructure cost optimization; and ability to build stable interfaces and collaborate with engineering and scientific stakeholders.
Desirable experience includes: implementing feature stores (Feast, Tecton, or internal systems); production work with Ray Serve, KServe, Triton, BentoML, SageMaker, Vertex AI, or custom gRPC inference services; serving and evaluating vision, document-understanding, embedding, or generative pipelines; expertise in GPU inference optimization, multi-model serving, edge inference, or Go/Rust performance-sensitive AI services; and background in regulated environments or open-source contributions to ML infrastructure projects.