SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fluidstack is building civilization-scale AI infrastructure, acquiring power, designing and operating data centers, and deploying compute at unprecedented speed and scale. The company's mission is to ensure frontier AI infrastructure is deployed in alignment with human freedom and choice.
As a Product Engineer on the Decision Team, you will own end-to-end systems that automate the delivery of gigawatts of compute. Your scope spans five major areas:
1. Fleet Health System: Build real-time telemetry and tiered health checks across Kubernetes and bare metal infrastructure, unified into a single trusted API. Design alarm correlation into incidents with drafted probable causes for on-call teams.
2. Repair & RMA Automation: Create a tracked workflow from failure detection through triage, parts sourcing, vendor returns, and return to service. Implement automatic failure routing, generate production engineer shift todo lists, and measure time-to-service recovery.
3. Hardware Qualification as Software: Compose burn-in, performance baselining, and validation into rack-level workflows. Enable thousands of accelerators to come online as repeatable runs with acceptance evidence attached.
4. Facility Operations System: Extend the maintenance system (lockout tagout, work orders) across multiple sites. Own the asset model, migration strategy, and cutover from legacy tools. Ensure asset registers load before external audits.
5. Runbook Automation: Convert SOPs, training records, and technician qualifications into auditable structured data. Build dashboards for site SLOs, deployment cycle time, and labor ramp reporting. Work forward-deployed beside production engineers and facility operators.
You will operate with full autonomy, own scope end-to-end, and drive everything forward with insane urgency. The company values first-principles reasoning, challenging assumptions, and shipping things that matter.
REQUIREMENTS:
- Shipped production code in Go, Python, or TypeScript; ability to pick up languages as needed
- Built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks
- Daily work with AI coding tools (Claude Code, Cursor) and autonomous agents
- Ability to identify problems, design solutions, and ship without waiting for direction
- Experience moving fast under deadline while building extensible foundations
- On-call rotation experience or close work with on-call engineers; track record of turning operational pain into quieter systems
- Product taste demonstrated in shipped work: obvious interfaces and workflows matching actual work patterns
- Bonus: Production engineering or SRE on large GPU fleets; hardware qualification or burn-in frameworks; BMC, Redfish, or IPMI tooling; CMMS, DCIM, or asset management systems; BMS/EPMS or SCADA; Prometheus and Grafana