SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sciforium is an AI infrastructure company building next-generation multimodal AI models and a proprietary high-efficiency serving platform. Backed by multi-million-dollar funding and direct AMD sponsorship, the company is scaling rapidly to develop the full stack powering frontier AI models and real-time applications.
As an AI Infrastructure Engineer, you will own the entire software stack of GPU clusters—from kernel tuning and GPU drivers through schedulers, containers, and ML frameworks. You'll define what production-ready nodes look like in software, authoring images, playbooks, and pipelines that transform freshly provisioned servers into fully validated GPU nodes. You'll maintain fleet consistency, upgradability, and performance while serving two demanding customer groups: foundation model training teams and model serving/product teams.
Key responsibilities include:
**OS Bring-Up & Node Lifecycle**: Own node software definition including versioned OS images, kernel tuning (NUMA, hugepages, IRQ affinity, cgroups), and GPU/NIC driver stacks. Build automated pipelines taking nodes from base OS to production-ready state.
**Validation & Fleet Management**: Create automated acceptance suites (DCGM diagnostics, NCCL/RCCL tests, bandwidth checks) that gate nodes before scheduler pool entry. Execute rolling kernel/driver/toolkit upgrades with minimal disruption and maintain driver-CUDA/ROCm-framework compatibility across the fleet.
**Infrastructure as Code**: Manage all node and cluster configuration through Ansible/SaltStack playbooks in Git with peer-reviewed changes, CI validation, and canary rollouts. Build provisioning pipelines (PXE, MaaS, Packer) ensuring reproducible node builds.
**Orchestration & Scheduling**: Deploy and operate GPU-enabled Kubernetes for inference workloads and Slurm for multi-node training. Maintain base images, registries, and container stacks for both NVIDIA and AMD accelerators.
**GPU Driver & ML Stack Engineering**: Build, deploy, and debug the full accelerator stack including CUDA toolkit, cuDNN, NCCL, ROCm, and RCCL. Maintain curated PyTorch and JAX environments and tune distributed performance across NVLink/NVSwitch and InfiniBand/RoCE fabrics.
**Advanced Debugging & Observability**: Own hard infrastructure problems including NCCL hangs, CUDA memory leaks, and throughput regressions. Implement software-layer monitoring (DCGM exporter, Prometheus/Grafana) and cluster efficiency reporting.
Required qualifications: 5+ years in systems/infrastructure engineering with significant GPU cluster, HPC, or large-scale ML infrastructure experience. Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or related field. Deep Linux internals expertise, hands-on experience with NVIDIA CUDA and/or AMD ROCm stacks on modern accelerators, production Kubernetes experience with GPU workloads, and working knowledge of HPC schedulers. Strong configuration management experience with Git-based workflows, provisioning/image tooling expertise, container fluency, and proficiency in Python and Bash. Working knowledge of NCCL, RDMA networking, and PyTorch/JAX runtime behavior required.
Nice-to-haves include direct experience supporting foundation model training teams, deploying inference/serving stacks (vLLM, Triton), and GPU/system profiling tools.