SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Multiverse Computing is a well-funded, fast-growing deep-tech company founded in 2019 and recognized by CB Insights as one of the 100 most promising AI companies globally. With 180+ employees, the company develops hyper-efficient software for quantum computing and AI applications.
You will own the deployment and integration strategy for Foundry, a sovereign AI development platform that must run across public cloud, private/sovereign cloud, on-premises Kubernetes, and air-gapped environments. This is a hands-on role where you will write production Python and Helm, and be the person who understands why deployments work.
Key responsibilities include:
- Design, build, and maintain Python services for integrating external applications, engines, and infrastructure into Foundry: manifest schemas, catalogs, versioning, install/upgrade/remove lifecycles, configuration, secrets handling, health reporting, and REST APIs.
- Define integration contracts for external engines to plug into Foundry's IAM, metadata layer, and UI; drive first integrations.
- Own Foundry's deployment packaging (Helm charts, operators, offline bundles) across Kubernetes distributions, cloud providers, sovereign clouds, and air-gapped environments.
- Integrate Foundry with customer infrastructure: OIDC/SAML identity providers, object storage, container registries, ingress, network policy, and GPU scheduling.
- Own CI/CD pipelines, release processes, and multi-environment integration test matrices.
- Instrument the platform with metrics, logs, and traces; write runbooks for production operations.
- Convert prototypes and one-off integrations into stable, documented, versioned services.
You will work within Platform Engineering alongside owners of metadata, catalog, and IAM modules, and with teams building compression, fine-tuning, serving, and orchestration capabilities.
REQUIREMENTS:
- 5+ years in platform, infrastructure, or backend engineering, with at least 2 years shipping and operating production workloads on Kubernetes.
- 3+ years writing production Python: owned backend services end-to-end (API design, data models, testing, packaging), not just automation scripts.
- Experience building and maintaining REST APIs in modern Python frameworks (FastAPI, Django REST Framework, or similar), including versioning, authentication, and input validation (Pydantic or equivalent).
- Experience with Python data access and migrations (SQLAlchemy, Alembic, or equivalent) on PostgreSQL, and schema design.
- Comfortable with async Python, structured logging, testable code with Pytest; use type hints and linters as standard practice.
- Experience interacting with Kubernetes and cloud APIs programmatically from Python (Kubernetes client, Boto3, or equivalent) for installers, controllers, or operational tooling.
- Deep, hands-on experience with Helm and Kubernetes packaging; have written and maintained charts others install.
- Experience deploying software into environments you do not control: on-premises, private cloud, or restricted-network/air-gapped installs.
- Experience with AWS: EKS, ECR, RDS, S3, Secrets Manager.
- Experience integrating with enterprise identity (OIDC, OAuth2, SAML) and managing secrets and configuration in production.
- Solid understanding of Docker, image building and hardening, and container registries.
- Proficiency with Git and CI/CD pipelines (GitLab CI or GitHub Actions).
- Product-oriented mindset: able to turn R&D scripts or prototypes into stable, usable services.
PREFERRED QUALIFICATIONS:
- Experience writing Kubernetes operators or controllers (Python with kopf, or Go), or GitOps tooling (ArgoCD, Flux).
- Experience building plugin, extension, or app-store style systems: manifest formats, lifecycle hooks, dependency and version resolution.
- Experience publishing Python packages or CLIs that others install (packaging, versioning, backwards compatibility).
- Experience with GPU workloads on Kubernetes (device plugins, node scheduling, NVIDIA GPU Operator) and LLM serving tools (vLLM, Triton, NIM).
- Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry) and exposing it for third-party components.
- Exposure to ML orchestration tooling (Flyte, Airflow, MLflow, SkyPilot).
- Go experience and track record contributing to open-source infrastructure projects.