SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Data Center Automation

Fluidstack - New York, NY, USA - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fluidstack is building civilization-scale AI infrastructure, acquiring power, designing and operating data centers at gigawatt scale. The company is singularly focused on deploying frontier compute infrastructure faster than anyone else, rethinking every layer of the stack from hardware to software. You will own the end-to-end design and delivery of critical systems that automate data center operations and fleet management. This is a high-impact role where you'll work directly with production engineers, facility operators, and hardware teams to turn operational pain into software systems. Key responsibilities include: - Build the fleet health system: real-time telemetry and tiered health checks across Kubernetes and bare metal infrastructure, rolled into a single trusted API with correlated alarms and incident routing to on-call with probable cause analysis. - Automate repair and RMA workflows: create a tracked system from failure detection through triage, parts ordering, vendor returns, and return to service, with automatic routing and generated work orders for production engineers. - Ship hardware qualification as software: compose burn-in, performance baselining, and validation into rack-level workflows so thousands of accelerators can be brought online repeatably with acceptance evidence attached. - Operate facility management systems: migrate legacy datacenter inventory and maintenance systems (lockout/tagout, work orders, asset registers) to a unified platform, managing the asset model and coordinating the switchover. - Turn runbooks into checked procedures: structure SOPs, training records, and technician qualifications as auditable data; build dashboards for site SLOs, deployment cycle time, and labor ramp reporting. You will work forward-deployed beside production engineers and facility operators, on-site and on rotation, building systems they use the next shift. The company operates with full autonomy, insane urgency, first-principles reasoning, and a focus on building something that matters. REQUIREMENTS: - Shipped production code in Go, Python, or TypeScript; ability to pick up whatever language the problem demands. - Built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks. - Daily work with AI coding tools like Claude Code and Cursor; experience getting agents to do useful work autonomously. - Ability to identify problems, design solutions, and ship without waiting for direction or approval. - Experience moving fast under deadline while building foundations other engineers can extend. - Sat on-call rotation or worked beside those who do; turned operational pain into systems that reduced pager burden. - Product taste demonstrated in shipped work: interfaces that feel obvious to operators, workflows matching actual work patterns. - Bonus: Production engineering or SRE on large GPU fleets; hardware qualification or burn-in frameworks; BMC, Redfish, or IPMI tooling; CMMS, DCIM, or asset management systems; BMS/EPMS or SCADA; Prometheus and Grafana.

Similar roles