SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Gallatin is rebuilding logistics infrastructure for U.S. national security missions. This MLOps role owns the complete lifecycle of AI/ML systems from training through production deployment, including infrastructure, release processes, evaluation, and observability.
You will build systems that move models from notebooks into operational environments—sometimes air-gapped racks in austere settings rather than cloud VPCs. The AI/ML team works on retrieval-grounded systems for doctrine and logistics, document/feature extraction, military symbol recognition, optimization models, and LLM agent platforms. Your infrastructure underpins all of it.
Key responsibilities:
Training & Serving Infrastructure: Own model training, fine-tuning, and batch inference across AWS (SageMaker, EKS) and on-premises GPU hardware. Stand up and optimize LLM inference serving (vLLM-class stacks, quantization, continuous batching, KV-cache sizing). Build for DDIL (local inference with configurable fallback and degraded-mode behavior) on hardware you don't control.
Release & Reproducibility: Build CI/CD for models and pipelines with versioned datasets, model registry, promotion gates, and reliable rollback. Own infrastructure-as-code, containerization, and GitOps deployment across dev clusters to disconnected enclaves. Make reproducibility a hard requirement—any customer-facing result must be reproducible from a commit and dataset version.
Evaluation & Observability: Build and own evaluation harnesses with regression suites for retrieval/extraction, LLM-as-judge pipelines with measured judge-human agreement, and adversarial/held-out test sets. Instrument production for drift, latency, cost, retrieval quality, and failure modes (including silent failures like fluent but incorrect answers). Make metrics defensible to external reviewers.
Data & Pipeline Ownership: Own ingestion, versioning, and lineage for logistics and doctrinal data from heterogeneous authoritative sources. Build embedding and feature pipelines, incremental indexing, and systems that keep retrieval indices fresh. Build human-in-the-loop infrastructure with confidence-scored routing, review queues, and feedback capture.
Secure & Accredited Deployment: Deploy and operate ML systems in IL5/IL6 environments, including air-gapped or restricted-network enclaves. Build release, observability, artifact-management, and incident-response workflows for those environments. Support ATO and continuous-authorization work with implementation evidence tied to security controls (NIST SP 800-171, NIST SP 800-53 Rev. 5, CMMC Level 2, FIPS 140-3, RMF/eMASS). Handle CUI and classified data correctly.
The team is small; you own domains end-to-end with no handoff. You'll work with seasoned entrepreneurs, operators, and technologists in a mission-driven environment where the stakes are real.
Requirements:
- 5+ years in MLOps, ML platform, or infrastructure engineering with meaningful time on systems with real users
- Strong Python skills and comfort in production codebases (not just notebooks)
- Deep Kubernetes and containerization experience plus infrastructure-as-code
- Production experience with AWS ML/Azure infrastructure (SageMaker, EKS, or equivalent)
- Hands-on GPU infrastructure experience: scheduling, utilization, memory sizing, cost
- Hands-on experience deploying and operating production software in IL5 or IL6 environments, including disconnected or restricted-network deployments
- Demonstrated shipping of LLM or ML systems to production and keeping them operational
- Experience building evaluation and monitoring for ML systems (not just vendor dashboards)
- Ability to reason about pipeline quality and identify when metrics measure the wrong thing
- Comfort with ambiguity and end-to-end domain ownership
- Willingness to learn the mission domain (logistics)
- U.S. citizenship required; ability to obtain and maintain U.S. government security clearance
Nice to have:
- Active clearance
- LLM serving and inference optimization (vLLM, TensorRT-LLM, quantization, prefix caching)
- Retrieval-grounded systems in production (hybrid retrieval, re-ranking, index freshness, citation quality)
- Edge, on-premises, or disconnected deployment experience
- ATO, continuous authorization, or classified environment operations support
- Defense, intelligence, or regulated environment background
- Palantir Foundry, PostgreSQL/pgvector, NATS/JetStream, ArgoCD
- CS, engineering, or related technical degree (or equivalent self-taught experience)