SlipstreamJobsFresh Startup & VC-Backed Jobs

Forward Deployed Engineer, Compute Operations

Fluidstack - New York, NY, USA - In-office - posted 2026-10-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fluidstack is building civilization-scale compute infrastructure for AI, acquiring power, designing and operating data centers with teams spanning hardware and software. The company is singularly focused on delivering 10-100s of GWs of compute faster than anyone else, rethinking every layer of the stack. As a Forward Deployed Engineer in Compute Operations, you will own end-to-end systems that keep Fluidstack's AI infrastructure running at scale. This is a high-autonomy role where you identify problems, design solutions, and ship them without waiting for direction. Key responsibilities include: - Build the fleet health system: real-time telemetry and tiered health checks on every machine across Kubernetes and bare metal, rolled into one API the whole company trusts, with alarms correlated into incidents that reach on-call with drafted probable cause. - Turn repair and RMA into generated work: one tracked flow from failure detection through triage, parts, vendor return, and return to service, where failure thresholds route machines to repair automatically and each production engineer's shift todo list is generated for them. - Ship hardware qualification as software: burn-in, performance baselining, and new hardware validation composed into rack-level workflows, so bringing thousands of accelerators online is repeatable and every machine enters production with acceptance evidence attached. - Run the facility on the same system as the fleet: migrate maintenance systems for lockout tagout and work orders across multiple sites, manage asset registers, and retire legacy datacenter inventory. - Turn every runbook into a checked procedure: SOPs, training records, and technician qualifications become structured data customers can audit, with site SLOs and deployment metrics reporting themselves on dashboards. You will work forward-deployed beside production engineers and facility operators, on site and on rotation. Fluidstack operates with full autonomy, insane urgency, first-principles reasoning, and a focus on building something that matters. The company is hiring people who care deeply about aligning frontier AI with human freedom. REQUIREMENTS: - Shipped production code in Go, Python, or TypeScript; ability to pick up whatever language the problem demands - Built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks - Work daily with AI coding tools like Claude Code and Cursor; experience getting agents doing useful work autonomously - Identify problems, design solutions, and ship without waiting for direction or approval - Moved fast under deadline while leaving foundations other engineers extended - Sat on-call rotation or worked beside those who do; turned operational pain into systems that made the pager quieter - Product taste demonstrated in shipped work: interfaces that are obvious to operators, workflows that match how work actually happens - Bonus: Production engineering or SRE on large GPU fleets; hardware qualification or burn-in frameworks; BMC, Redfish, or IPMI tooling; CMMS, DCIM, or asset management systems; BMS/EPMS or SCADA; Prometheus and Grafana

Similar roles