SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Workist is building AI agents, and this role owns the platform infrastructure they run on. You will be responsible for two cloud environments (Azure primary, AWS secondary), Kubernetes clusters across production, staging, and ML workloads, and the entire CI/CD pipeline from merge request to production. You set your own quarterly roadmap for platform evolution, propose infrastructure themes, and deliver them while handling ad-hoc incidents, security findings, and developer support—roughly one-third of your time.
Key responsibilities include:
**Own the Platform**: Design and maintain cloud architecture across Azure and AWS (subscriptions, networks, identities, environment isolation). Deliver platform capabilities the product needs next, from GPU capacity and LLM gateways to new environments. Set standards for Kubernetes operations and managed data services, ensuring capacity planning, safe rollouts, and cost efficiency.
**Keep It Quiet**: Own end-to-end security posture including least-privilege access, network isolation, scanner findings, and pentest remediation. Manage a self-running patch cadence and serve as technical counterpart for audits. Own observability and alerting so every alert is actionable. Lead platform incident response and quarterly backup/disaster-recovery tests.
**Make the Team Faster**: Own the merge-to-production path with fast builds, reliable deployments, and on-demand environments. Build infrastructure for AI models including LLM gateways, region failover, and GPU capacity planning.
**Technical Stack**: Kubernetes (Azure AKS), Helm, GitOps/Flux, Terraform, GitLab CI/CD, managed services (Postgres, OpenSearch, Redis, blob storage), Python application debugging.
**Recent Examples**: Debugged a Redis instance silently dropping TCP connections that crashed Celery workers. Implemented region failover and load balancing for Azure OpenAI deployments. Migrated clusters from nginx ingress to Kubernetes Gateway API.
You report to the CTO and are the voice of the platform in engineering decisions. The role is individual-contributor focused—you own the platform end-to-end rather than managing a team.
**Requirements**:
- Years of production infrastructure experience, ideally as the sole owner rather than one of many in a large ops team
- Kubernetes and Terraform in production: cluster upgrades, Helm, GitOps, infrastructure-as-code for entire cloud estates
- Production experience on Azure or AWS (Azure preferred; AWS background acceptable if willing to go deep on Azure)
- Active Python skills: read, debug, and fix Python codebases; fix application-layer issues when needed
- Security as a habit: patch management, pentest findings, least-privilege access design, pragmatic prioritization
- Clear writing: runbooks, incident write-ups, architecture decisions that others can follow
- Fluent English (German not required but helpful for customer/vendor conversations)