SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
HUD is building infrastructure for reinforcement learning training data and evaluations for frontier AI agents, plus a marketplace to sell these services to frontier labs. The platform is used by frontier labs, Fortune 500 companies, and startups. The company has raised $16M from top VCs and was part of Y Combinator W25.
You will own the reliability, scale, performance, and developer experience of HUD's core infrastructure and backend systems. This is not a pure infrastructure role—the ideal candidate combines strong production infrastructure experience with backend engineering judgment. You'll work across AWS, Kubernetes, Terraform, CI/CD, observability, and backend services to make HUD faster, more reliable, cheaper to run, and easier for engineers to build on.
Key responsibilities include:
- Owning production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services
- Building and maintaining AWS infrastructure with Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management
- Designing and improving backend and platform systems for scale, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths
- Defining and improving dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows
- Building reliable CI/CD, release automation, environment management, and deployment workflows
- Writing clean, maintainable code to automate systems, improve backend services, and create internal tooling
The team is ~25 people, mostly full-time in-person but with some remote flexibility. The team includes 4 International Olympiad medalists, serial AI startup founders, and researchers with publications at ICLR and NeurIPS. The company has 8 figures in funding and is scaling profitably to meet strong demand.
REQUIREMENTS:
- Must have owned production cloud infrastructure for a high-availability, user-facing platform with responsibility for uptime, performance, deployment safety, and cost
- Deep experience with AWS infrastructure and containerized systems; strong preference for Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, load balancers, networking, and secrets management
- Experience building or operating CI/CD, environment management, release automation, observability, alerting, and incident response systems
- Strong backend engineering judgment and ability to reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes
- Ability to write clean, maintainable code and apply strong software engineering judgment across product architecture, infrastructure, backend systems, and developer workflows
Strong candidates may also have:
- Experience operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms
- Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services
- Experience reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability
- Experience building internal platforms or tools that make engineers faster without hiding too much complexity
The company prioritizes technical aptitude, ownership, and learning potential over years of experience. Full-time employment. Offices in San Francisco or Singapore; open to remote candidates who can work 70-80% overlap with either timezone. Visa sponsorship provided for strong candidates to US or Singapore.