SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: SGD 200,000 - 400,000 / annual
Inferact, founded by the creators and core maintainers of vLLM, is building the world's AI inference engine. The company sits at the intersection of models and hardware, focused on making inference cheaper and faster.
You will own and operate Inferact's high-performance GPU compute infrastructure across neo-cloud and dedicated compute providers. This is a hands-on cluster administration role with end-to-end responsibility for cluster health, GPU availability, monitoring, alerting, scheduling, access control, diagnostics, and incident response. You'll ensure infrastructure is healthy, available, observable, and usable around the clock, directly supporting engineering teams building and testing vLLM.
Key responsibilities include:
- Taking ownership of GPU cluster operations across multiple providers
- Managing cluster health, GPU availability, and resource scheduling
- Implementing monitoring, alerting, and incident response workflows
- Working with engineering leadership to standardize provisioning, debugging, and scaling practices
- Automating operational workflows to reduce toil
- Diagnosing and resolving urgent infrastructure incidents that block engineering teams
- Managing GPU server operations including driver management, health monitoring, and hardware diagnostics
You'll work closely with infrastructure owners and engineering leadership to ensure compute infrastructure scales efficiently and reliably as Inferact grows.
REQUIREMENTS:
Minimum qualifications:
- Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar
- Hands-on experience administering large compute clusters, HPC environments, university/research clusters, supercomputing systems, or production GPU clusters
- Strong Linux systems administration fundamentals: networking, processes, storage, package management, shell scripting, logs, access control, system debugging
- Experience operating GPU servers: driver management, GPU health monitoring, node failures, memory errors, scheduler issues, hardware diagnostics
- Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent
- Ability to own urgent infrastructure incidents end-to-end
- Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar
Preferred qualifications:
- Experience operating GPU compute across providers (Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar)
- Experience improving cluster utilization and debugging scheduling/resource contention issues
- Familiarity with high-performance GPU networking (InfiniBand, RoCE, NVLink/NVSwitch, RDMA, NCCL)
- Experience with HPC/ML storage systems (NFS, Lustre, Ceph, distributed filesystems)
- Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, infrastructure security
- Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations
- GPU/HPC infrastructure management in university labs, national labs, research institutions, AI infrastructure companies, hedge funds, HFT firms, or large-scale ML platforms
- Built monitoring, alerting, runbooks, health checks, or remediation workflows that reduced operational toil
- Operated Kubernetes clusters for ML/GPU workloads at scale
- Standardized provisioning, diagnostics, monitoring across multiple compute providers
- Carried operational responsibility for infrastructure used by many engineers/researchers