SlipstreamJobsFresh Startup & VC-Backed Jobs

Distributed Systems ML Infrastructure Engineer

OpenTeams - Washington, DC, United States - Hybrid - posted 2026-09-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 145,000 - 250,000 / annual

OpenTeams, founded by Travis Oliphant (creator of NumPy and SciPy) and built by veterans of the open-source ecosystem (NumPy, SciPy, PyTorch, Jupyter), is building an AI platform that enables enterprises and governments to own, control, and govern their own AI rather than renting intelligence from vendors. You will design and implement the core distributed services of a containerized, API-first AI platform. This is a hands-on engineering role focused on building the platform infrastructure itself, not integrating third-party solutions. Key responsibilities include: - Design and implement platform services for workflow orchestration, data ingestion, results management, and model serving, exposed through documented APIs - Build and operate model gateway and serving services with policy enforcement, usage accounting, and audit logging - Maintain provider abstraction layers so the platform runs on AWS-native managed services while remaining deployable across other cloud and dedicated environments - Deploy platform releases into classified environments, perform data source integration, and execute validation procedures - Verify environment parity and validate platform performance against documented workload models through load testing - Constrain platform dependencies to services available in target environments and gate development-only dependencies behind feature flags - Reproduce and fix high-side defects through sanitized feedback paths You'll work primarily on unrestricted infrastructure with an open-source toolchain. Engineers with appropriate access also carry releases into controlled production environments and validate the platform in place, providing visibility into real-world deployment. Required qualifications: 6+ years in distributed systems, platform engineering, or infrastructure engineering; production experience with Kubernetes and containerized workloads on major cloud platforms (especially Amazon EKS); experience supporting ML workloads in production (model serving, GPU scheduling, large-scale data pipelines); proficiency in Python, Go, or comparable language; experience with infrastructure-as-code tools (Terraform, Helm); API-first service design; and capacity/performance testing. U.S. citizenship and security clearance eligibility required. Active TS/SCI clearance with CI polygraph strongly preferred.

Similar roles