SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior/Staff Kubernetes Infrastructure Engineer

Fal - Remote - Remote - posted 2026-08-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fal is a generative media infrastructure platform enabling developers and enterprises to move AI products from idea to production at scale. The company provides unified infrastructure for high-performance inference, orchestration, and observability. You will architect and operate the high-performance compute environments delivered to customers, spanning bare-metal servers, GPU-enabled VMs, Kubernetes clusters, Slurm clusters, distributed storage, and high-speed networking. This is a full-stack infrastructure role requiring deep expertise across the entire compute lifecycle. Key responsibilities include: designing and automating the complete lifecycle of customer compute environments from provisioning through upgrades and decommissioning; provisioning dedicated Kubernetes and Slurm clusters tailored to specific workloads; building and maintaining Linux images with automated OS-provisioning workflows; operating the NVIDIA GPU stack including drivers, GPU Operator, container toolkit, device plugins, and monitoring; designing Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP; configuring distributed and shared storage for high-performance workloads; building monitoring, alerting, diagnostics, and automated recovery systems; developing reusable tooling, standards, documentation, and runbooks; and collaborating with customers and internal teams to translate workload requirements into infrastructure designs. Required qualifications: 5+ years building and operating production Linux infrastructure; strong production experience with Kubernetes on bare metal including bootstrapping, HA control planes, etcd, containerd, CNI, CSI, and observability; Linux virtualization experience with KVM/QEMU, libvirt, and VFIO device passthrough; NVIDIA GPU operations on Linux and Kubernetes; strong networking fundamentals including TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting; practical scripting and configuration-management tools like Ansible; ability to diagnose complex cross-layer infrastructure issues; and strong communication skills. Nice-to-have skills include production Slurm experience, high-performance networking (NVLink, InfiniBand, RoCEv2, GPUDirect RDMA), NUMA/CPU pinning, SR-IOV/DPDK, distributed storage systems (Ceph, Lustre, Weka), KubeVirt/OpenStack, network security (IPsec, WireGuard), bare-metal management tools, network automation platforms, AI training/inference infrastructure experience, and Python or Go proficiency.

Similar roles