SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 145,000 - 250,000 / annual
OpenTeams, founded by Travis Oliphant (creator of NumPy and SciPy) and built by veterans of the open-source ecosystem (NumPy, SciPy, PyTorch, Jupyter), is seeking a Senior Infrastructure Engineer to build and operate a Kubernetes-based platform supporting secure AI test and evaluation capabilities for Government teams.
You will own the platform underlying AI system assessment, supporting demanding AI/ML workloads including GPU scheduling, large-scale data movement, reproducible test execution, and multi-tenant isolation. The platform must operate reliably in restricted Government environments with limited connectivity, no assumption of outbound internet access, and no managed cloud services.
Key responsibilities include:
- Building and operating a Kubernetes platform for AI test and evaluation frameworks
- Implementing GPU scheduling, workload orchestration, resource management, and multi-tenant isolation
- Designing infrastructure-as-code, GitOps workflows, and automated deployment pipelines using OpenTofu, Terraform, Helm, Argo CD, and Kubernetes operators
- Developing reusable, modular infrastructure components that can be independently deployed and operated
- Contributing to Nebari and other open-source Kubernetes and MLOps projects
- Owning platform reliability: capacity planning, upgrade strategies, failure-mode analysis, backup/recovery, and operational readiness
- Designing and implementing observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads
- Developing operational runbooks and documentation for deployment, operation, and troubleshooting
- Deploying and hardening infrastructure in secure, disconnected, or limited-connectivity Government environments
- Supporting security authorization and compliance through documentation, hardened configurations, and repeatable deployment processes
- Integrating automated security tooling for container scanning, SAST/DAST, artifact signing, and policy enforcement
- Collaborating with Government stakeholders, security personnel, software engineers, and ML engineers
- Providing technical leadership, contributing to engineering standards, and mentoring junior team members
- Working effectively in a remote, distributed team using asynchronous communication
Required qualifications:
- U.S. citizenship and ability to obtain and maintain a Secret security clearance
- 6+ years of hands-on infrastructure, platform, DevOps, or SRE experience supporting production systems
- Strong understanding of infrastructure engineering principles: scalability, reliability, observability, security, automation
- Production experience with Kubernetes including workload scheduling, resource management, multi-tenant environments
- Experience with automated security tooling (container scanning, SAST/DAST, artifact signing, policy enforcement)
- Proficiency with infrastructure-as-code tools (Terraform, OpenTofu, Pulumi, or equivalent)
- Experience with at least one major cloud platform (AWS, Azure, Google Cloud) including networking, security, storage, compute
- Experience implementing monitoring and observability solutions
This is a fully remote, U.S.-based role working with a distributed team that relies heavily on asynchronous communication.