SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Shield AI is a venture-backed defense-tech company developing autonomous systems, aircraft, and simulation technologies to protect service members and civilians. The company operates globally with offices across the U.S., Europe, the Middle East, and Asia-Pacific.
You will own the technical strategy, architecture, and roadmap for automated testing, verification, and MLOps quality across the Hivemind autonomy ecosystem. This is a company-wide technical leadership role for test infrastructure and quality assurance.
Key responsibilities include:
- Design and maintain scalable test frameworks for autonomy software, backend services, APIs, operator-facing applications, and distributed hardware environments
- Lead functional, integration, regression, system, performance, reliability, and end-to-end testing across simulation, edge-compute, software-in-the-loop, and hardware-in-the-loop environments
- Build automated ML validation pipelines covering data quality, training reproducibility, model accuracy, robustness, regression, latency, resource utilization, and system integration
- Establish CI/CD and continuous training workflows with versioning and traceability for datasets, models, configurations, evaluation results, and deployment artifacts
- Develop scenario-based validation for autonomy models, including edge cases, degraded sensing or communications, distribution shifts, and representative mission conditions
- Create observability, analytics, and failure-triage capabilities for software behavior, model and data drift, inference health, test results, and production performance
- Build Python automation that improves test execution, parallelization, reporting, environment setup, experiment comparison, and developer productivity
- Create test harnesses, simulators, stubs, mocks, and synthetic data capabilities
- Collaborate with software, autonomy, machine learning, data, simulation, and systems engineers to define verification strategies
- Develop and govern AI-assisted engineering workflows using coding agents and LLM-based tools for test generation, log analysis, debugging, and failure triage while maintaining security, reproducibility, and traceability
Requirements:
- Typically 8+ years of relevant experience in software engineering, test infrastructure, developer tooling, MLOps, systems integration, or systems verification, or equivalent combination of experience and demonstrated impact
- 5+ years of experience building scalable automation frameworks or developer tooling in Python
- Demonstrated success designing test, CI/CD, or MLOps infrastructure used across multiple engineering teams
- Experience validating machine learning systems across the data, training, evaluation, packaging, deployment, and monitoring lifecycle
- Understanding of ML quality risks such as data leakage, training-serving skew, nondeterminism, distribution shift, drift, model regression, and statistical acceptance criteria
- Experience defining model-performance baselines, automated evaluation suites, release thresholds, and candidate-to-production comparison workflows
- Experience testing GPU-accelerated infrastructure and workloads, including GPU scheduling, allocation, utilization, and resource contention in Kubernetes environments
- Experience with performance benchmarking, profiling, and observability for GPU workloads, including identifying compute, memory, storage, networking, and data-loading bottlenecks
- Experience validating multi-tenant Kubernetes environments, including RBAC, resource quotas, workload isolation, and scheduling behavior
- Experience qualifying integrated hardware and software systems, including automated validation of compute, GPU, storage, networking, drivers, firmware, and deployed software configurations
- Strong system-design skills and experience testing distributed systems, backend services, APIs, and integrated hardware and software environments
- Experience developing integration and regression strategies for internally developed, third-party, open-source, and partner software, including dependency management, compatibility testing, and upgrades
- Strong understanding of asynchronous and concurrent Python programming for scalable automation and parallel test execution
- Experience with package and dependency management, reproducible environments, and build systems such as Conan, pip, setuptools, Poetry, Nix, or similar
- Experience with automated observability, log collection, analytics, reporting, and root-cause analysis in complex software, data, and infrastructure systems
- Experience working in Linux-based development environments
Preferred qualifications include: model registries and experiment tracking (MLflow, Kubeflow, Weights & Biases, SageMaker, Vertex AI), GPU scheduling platforms (Run:ai, NVIDIA GPU Operator, KAI Scheduler, Kueue, Volcano), NVIDIA GPU infrastructure expertise, GPU profiling tools, autonomy model validation experience, embedded/edge-compute testing, production server qualification, reliability and fault-injection testing, reproducible installation and rollback validation, containers and Kubernetes, Go or TypeScript proficiency, Python-to-C/C++ interoperability, aerospace/robotics/autonomy/embedded systems background, and familiarity with software-in-the-loop, hardware-in-the-loop, requirements-based verification, and standards such as DO-178C and MIL-STD-882.