SlipstreamJobsFresh Startup & VC-Backed Jobs

Platform Engineer, AI/ML Infrastructure

OpenTeams - Remote - Hybrid - posted 2026-09-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 120,000 - 250,000 / annual

OpenTeams, founded by Travis Oliphant (creator of NumPy and SciPy) and built by engineers from the open-source ecosystem (NumPy, SciPy, PyTorch, Jupyter), is hiring Platform Engineers to build AI infrastructure that enterprises and governments own outright—including the infrastructure, data, models, and evidence of system behavior. The role spans the full depth of an AI platform: Kubernetes clusters scheduling GPU workloads, multi-tenant isolation, platform services (workflow orchestration, data ingest, model serving, policy enforcement, audit logging), CI/CD delivery pipelines, and cloud infrastructure. A key constraint is that much of this must run in environments without normal cloud resources or internet connectivity—making portability, reproducibility, and operability design inputs from day one. The team builds on open source (Kubernetes, Terraform, OpenTofu, Argo, Prometheus, Nebari) and contributes upstream. This posting covers multiple roles spanning mid-level through senior; the company assigns levels based on actual experience rather than years alone. Key Responsibilities: - Build and operate Kubernetes-based infrastructure for AI/ML workloads, including GPU scheduling, resource management, and multi-tenant isolation - Design and implement platform services for orchestration, data ingest, model serving, and results management behind documented APIs - Write infrastructure as code and build GitOps pipelines for reproducible environments - Build and operate CI/CD pipelines producing versioned, signed, scanned release artifacts with deployment documentation - Own reliability: capacity planning, upgrade paths, failure-mode analysis, backup/recovery, incident response, postmortems - Implement monitoring, logging, tracing, alerting, and define service level objectives - Deploy and validate the platform in restricted, disconnected, or limited-connectivity environments - Keep the platform portable by constraining dependencies to confirmed available resources - Write runbooks and operational documentation executable by other engineers - Contribute to Nebari and other open-source infrastructure and MLOps projects - Work with security engineers, government stakeholders, and other engineers to turn requirements into resilient systems - Collaborate asynchronously across a distributed team Two positions available: one hybrid (Washington DC, Denver CO, or Colorado Springs CO; requires active TS/SCI clearance; up to 15% travel); one fully remote US (clearance not required but must be willing to obtain). Requirements: - U.S. citizenship and ability to obtain and maintain a U.S. security clearance - Four or more years of hands-on experience building or operating production infrastructure, platforms, or distributed systems - Production experience with Kubernetes and containerized workloads - Experience with at least one major cloud platform (AWS, Azure, or Google Cloud) - Experience with infrastructure as code and CI/CD (Terraform, OpenTofu, Pulumi, Helm, or similar) - Working proficiency in Python, Go, Bash, or comparable language - Experience implementing or operating production monitoring and observability - Ability to write documentation, runbooks, and deployment procedures others can follow - Ability to work independently and collaborate in a remote, distributed team Nice to Have (not required): - Active U.S. security clearance, particularly TS/SCI with CI polygraph - Experience deploying software in air-gapped, disconnected, or restricted environments - Experience with DoD, Intelligence Community, or comparably regulated programs - Familiarity with Risk Management Framework, NIST 800-53/800-171, or similar frameworks - Experience supporting Authorization to Operate or continuous ATO models - Supply chain security work (hardened images, artifact signing, SBOM generation, scanning, policy enforcement) - Familiarity with cross-domain solutions, guards, data diodes, or transfer mechanisms - Experience with classified cloud environments (AWS Secret/Top Secret regions) - DoD 8140/8570 qualifying certification (Security+, CISSP, CASP+, CISM) or willingness to obtain - Experience building MLOps pipelines or AI/ML infrastructure - Experience with GPU scheduling, distributed inference, or large-scale data/evaluation pipelines - Experience with model-serving or gateway frameworks (KServe, vLLM, LLM-D) - Experience designing API-first, vendor-agnostic platforms across multiple environments - Experience with agentic workflow frameworks or multi-step AI pipeline orchestration - Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects; Nebari experience - Familiarity with data sovereignty and privacy requirements for enterprise/government AI systems - Experience leading technical initiatives, setting engineering standards, or mentoring engineers - Experience supporting rapid prototyping or defense innovation initiatives

Similar roles