SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Compute Operations

Fluidstack - New York, NY, USA - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fluidstack is building civilization-scale compute infrastructure for AI at unprecedented speed and scale. The company acquires power, designs and operates data centers, and is rethinking every layer of the stack to deliver 10-100s of GWs of compute faster than anyone else. As a Software Engineer in Compute Operations, you will own end-to-end systems that keep Fluidstack's massive GPU fleet healthy and operational. You'll work on several interconnected challenges: **Fleet Health & Monitoring**: Build real-time telemetry and tiered health checks across Kubernetes and bare-metal infrastructure, unified into a single trusted API. Design alarm correlation systems that surface incidents to on-call with probable cause analysis. **Repair & RMA Automation**: Create a tracked workflow from failure detection through triage, parts sourcing, vendor returns, and return to service. Implement automatic failure routing and generate production engineer shift todo lists from live data. **Hardware Qualification**: Compose burn-in, performance baselining, and validation into rack-level workflows so thousands of accelerators can be brought online repeatably, with acceptance evidence attached to each machine. **Facility Operations**: Extend maintenance systems (lockout/tagout, work orders) across multiple sites. Own the asset model, data migration, and cutover from legacy tools. Build audit-ready structured data for SOPs, training records, and technician qualifications. **Operational Excellence**: Work forward-deployed beside production engineers and facility operators. Turn runbooks into checked procedures. Build dashboards for site SLOs, deployment cycle time, and labor ramp reporting. You'll operate with full autonomy, own scope end-to-end, and move with insane urgency. The company values first-principles thinking, challenging assumptions, and shipping things that actually matter. **Requirements:** - Shipped production code in Go, Python, or TypeScript; ability to pick up languages as needed - Built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks - Daily experience with AI coding tools (Claude Code, Cursor) and autonomous agents - Ability to identify problems, design solutions, and ship without waiting for direction - Track record of moving fast under deadline while building extensible foundations - Experience on-call or working closely with on-call engineers; ability to turn operational pain into systems - Strong product taste; interfaces and workflows that match how work actually happens - Bonus: Production engineering or SRE on large GPU fleets; hardware qualification or burn-in frameworks; BMC/Redfish/IPMI tooling; CMMS/DCIM/asset management; BMS/EPMS or SCADA; Prometheus and Grafana

Similar roles