SlipstreamJobsFresh Startup & VC-Backed Jobs

Product Engineer, Compute Operations

Fluidstack - New York, NY, USA - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fluidstack is building civilization-scale compute infrastructure for AI, acquiring power, designing and operating data centers, and rethinking every layer of the stack. The company is singularly focused on delivering 10–100s of GWs of compute faster than anyone else, with teams spanning hardware and software. As a Product Engineer in Compute Operations, you will own end-to-end systems that keep Fluidstack's massive GPU fleet healthy and operational. Your scope includes: **Fleet Health & Monitoring**: Build real-time telemetry and tiered health checks across Kubernetes and bare-metal machines, unified into a single trusted API. Correlate alarms into incidents with drafted probable causes for on-call teams. **Repair & RMA Automation**: Transform repair workflows into generated work—from failure detection through triage, parts ordering, vendor returns, and return to service. Failure thresholds automatically route machines; production engineers receive auto-generated shift todo lists; the system reports time-to-service metrics. **Hardware Qualification as Software**: Compose burn-in, performance baselining, and validation into rack-level workflows so thousands of accelerators come online repeatably, with acceptance evidence attached in a knowledge graph. **Facility Operations**: Extend the fleet maintenance system (lockout/tagout, work orders) across multiple sites. Own the asset model, migration strategy, and cutover from legacy tools. Ensure asset registers load before external audits. **Runbook Automation**: Convert SOPs, training records, and technician qualifications into auditable structured data. Build dashboards for site SLOs, deployment cycle time, and labor ramp. You'll work forward-deployed on-site and on rotation with production engineers and facility operators. The company operates with full autonomy, insane urgency, first-principles reasoning, and a focus on building something that matters. You'll identify problems, design solutions, and ship without waiting for direction. **Requirements:** - Shipped production code in Go, Python, or TypeScript; ability to pick up whatever language the problem demands - Built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks - Daily use of AI coding tools (Claude Code, Cursor) and experience getting agents to do useful autonomous work - Ability to identify problems, design solutions, and ship independently - Experience moving fast under deadline while building foundations others can extend - Sat on-call rotation or worked closely with on-call engineers; turned operational pain into quieter systems - Strong product taste: interfaces and workflows that match how work actually happens - Bonus: Production engineering or SRE on large GPU fleets; hardware qualification or burn-in frameworks; BMC, Redfish, or IPMI tooling; CMMS, DCIM, or asset management systems; BMS/EPMS or SCADA; Prometheus and Grafana

Similar roles