SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Infrastructure Engineer - AI/ML Platform

OpenTeams - Remote - Remote - posted 2026-09-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 145,000 - 250,000 / annual

OpenTeams, founded by Travis Oliphant (creator of NumPy and SciPy) and built by veterans of the open-source ecosystem (NumPy, SciPy, PyTorch, Jupyter), is seeking a Senior Infrastructure Engineer to build and operate a Kubernetes-based platform supporting secure AI test and evaluation capabilities for Government teams. You will own the platform underlying AI system assessment, supporting demanding AI/ML workloads including GPU scheduling, large-scale data movement, reproducible test execution, and multi-tenant isolation. The platform must operate reliably in restricted Government environments with limited connectivity, no assumption of outbound internet access, and no managed cloud services. Key responsibilities include: - Building and operating a Kubernetes platform for AI test and evaluation frameworks - Implementing GPU scheduling, workload orchestration, resource management, and multi-tenant isolation - Designing infrastructure-as-code, GitOps workflows, and automated deployment pipelines using OpenTofu, Terraform, Helm, Argo CD, and Kubernetes operators - Developing reusable, modular infrastructure components that can be independently deployed and operated - Contributing to Nebari and other open-source Kubernetes and MLOps projects - Owning platform reliability: capacity planning, upgrade strategies, failure-mode analysis, backup/recovery, and operational readiness - Designing and implementing observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads - Developing operational runbooks and documentation for deployment, operation, and troubleshooting - Deploying and hardening infrastructure in secure, disconnected, or limited-connectivity Government environments - Supporting security authorization and compliance through documentation, hardened configurations, and repeatable deployment processes - Integrating automated security tooling for container scanning, SAST/DAST, artifact signing, and policy enforcement - Collaborating with Government stakeholders, security personnel, software engineers, and ML engineers - Providing technical leadership, contributing to engineering standards, and mentoring junior team members - Working effectively in a remote, distributed team using asynchronous communication Required qualifications: - U.S. citizenship and ability to obtain and maintain a Secret security clearance - 6+ years of hands-on infrastructure, platform, DevOps, or SRE experience supporting production systems - Strong understanding of infrastructure engineering principles: scalability, reliability, observability, security, automation - Production experience with Kubernetes including workload scheduling, resource management, multi-tenant environments - Experience with automated security tooling (container scanning, SAST/DAST, artifact signing, policy enforcement) - Proficiency with infrastructure-as-code tools (Terraform, OpenTofu, Pulumi, or equivalent) - Experience with at least one major cloud platform (AWS, Azure, Google Cloud) including networking, security, storage, compute - Experience implementing monitoring and observability solutions This is a fully remote, U.S.-based role working with a distributed team that relies heavily on asynchronous communication.

Similar roles