SlipstreamJobsFresh Startup & VC-Backed Jobs

Platform Engineer

Multiverse Computing - Remote - Hybrid

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Multiverse Computing is a well-funded, fast-growing deep-tech company founded in 2019 and recognized by CB Insights as one of the 100 most promising AI companies globally. With 180+ employees, the company develops hyper-efficient software for quantum computing and AI applications. You will own the deployment and integration strategy for Foundry, a sovereign AI development platform that must run across public cloud, private/sovereign cloud, on-premises Kubernetes, and air-gapped environments. This is a hands-on role where you will write production Python and Helm, and be the person who understands why deployments work. Key responsibilities include: - Design, build, and maintain Python services for integrating external applications, engines, and infrastructure into Foundry: manifest schemas, catalogs, versioning, install/upgrade/remove lifecycles, configuration, secrets handling, health reporting, and REST APIs. - Define integration contracts for external engines to plug into Foundry's IAM, metadata layer, and UI; drive first integrations. - Own Foundry's deployment packaging (Helm charts, operators, offline bundles) across Kubernetes distributions, cloud providers, sovereign clouds, and air-gapped environments. - Integrate Foundry with customer infrastructure: OIDC/SAML identity providers, object storage, container registries, ingress, network policy, and GPU scheduling. - Own CI/CD pipelines, release processes, and multi-environment integration test matrices. - Instrument the platform with metrics, logs, and traces; write runbooks for production operations. - Convert prototypes and one-off integrations into stable, documented, versioned services. You will work within Platform Engineering alongside owners of metadata, catalog, and IAM modules, and with teams building compression, fine-tuning, serving, and orchestration capabilities. REQUIREMENTS: - 5+ years in platform, infrastructure, or backend engineering, with at least 2 years shipping and operating production workloads on Kubernetes. - 3+ years writing production Python: owned backend services end-to-end (API design, data models, testing, packaging), not just automation scripts. - Experience building and maintaining REST APIs in modern Python frameworks (FastAPI, Django REST Framework, or similar), including versioning, authentication, and input validation (Pydantic or equivalent). - Experience with Python data access and migrations (SQLAlchemy, Alembic, or equivalent) on PostgreSQL, and schema design. - Comfortable with async Python, structured logging, testable code with Pytest; use type hints and linters as standard practice. - Experience interacting with Kubernetes and cloud APIs programmatically from Python (Kubernetes client, Boto3, or equivalent) for installers, controllers, or operational tooling. - Deep, hands-on experience with Helm and Kubernetes packaging; have written and maintained charts others install. - Experience deploying software into environments you do not control: on-premises, private cloud, or restricted-network/air-gapped installs. - Experience with AWS: EKS, ECR, RDS, S3, Secrets Manager. - Experience integrating with enterprise identity (OIDC, OAuth2, SAML) and managing secrets and configuration in production. - Solid understanding of Docker, image building and hardening, and container registries. - Proficiency with Git and CI/CD pipelines (GitLab CI or GitHub Actions). - Product-oriented mindset: able to turn R&D scripts or prototypes into stable, usable services. PREFERRED QUALIFICATIONS: - Experience writing Kubernetes operators or controllers (Python with kopf, or Go), or GitOps tooling (ArgoCD, Flux). - Experience building plugin, extension, or app-store style systems: manifest formats, lifecycle hooks, dependency and version resolution. - Experience publishing Python packages or CLIs that others install (packaging, versioning, backwards compatibility). - Experience with GPU workloads on Kubernetes (device plugins, node scheduling, NVIDIA GPU Operator) and LLM serving tools (vLLM, Triton, NIM). - Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry) and exposing it for third-party components. - Exposure to ML orchestration tooling (Flyte, Airflow, MLflow, SkyPilot). - Go experience and track record contributing to open-source infrastructure projects.

Similar roles