SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer

Crusoe - Dublin, Ireland - In-office - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Crusoe is building the world's only vertically integrated AI infrastructure company, owning and operating every layer of the stack from electrons to tokens. The company is solving the AI compute energy bottleneck with an energy-first approach to infrastructure. As a Staff Software Engineer on the Cloud Availability Platform team, you will help build Conductor—a self-driving control plane designed to predict, decide, and act autonomously to keep tens of thousands of accelerators running at peak efficiency. You will have true ownership over greenfield services, building distributed systems that optimize power, cost, and useful compute in real time. Working alongside staff and principal engineers, you will tackle rare infrastructure challenges such as energy-aware compute scheduling, closed-loop remediation, and failure prediction on noisy hardware telemetry. Key responsibilities include: - Design, build, and operate greenfield distributed services, control planes, and closed-loop remediation systems that automatically drain, checkpoint, replace, and resume live AI workloads without manual intervention - Partner closely with staff and principal engineers to review architectures, establish safety rails for autonomous operations, and share technical knowledge across the cloud services team - Collaborate daily with cross-functional partners spanning hardware engineering, energy management, data center construction, and customer support to align system design with physical infrastructure - Own development of core platform features such as unified observability pipelines, energy-aware compute scheduling, straggler detection, and self-qualifying hardware pipelines to maximize overall system goodput This is a full-time position for a problem-solving engineer who thrives in ambiguity and has a strong bias toward shipping. You will transform complex operational signals into a single, programmable control plane that turns tens of thousands of accelerators across sites into one cohesive, logical system. REQUIREMENTS: - Minimum 2+ years of experience building, shipping, and operating backend or infrastructure systems in production (e.g., distributed services, control planes, schedulers, or observability pipelines) - Demonstrated mastery of state reconciliation, retries and idempotency, consistency tradeoffs, and autonomous automation on live infrastructure - Strong engineering fundamentals using a modern systems language such as Go, Rust, or C++ - High comfort level working through ambiguous problem statements, proposing clear designs, and driving features from concept to production - Strong bias toward action and iterative delivery, favoring practical production software over over-engineered documentation - Bachelor's degree or equivalent practical experience in Computer Science, Engineering, or a related technical discipline BONUS QUALIFICATIONS: - Hands-on experience or deep familiarity with GPU health telemetry, NVLink/InfiniBand/RoCE networking fabrics, or hardware thermal and power behavior - Prior experience designing scalable observability and telemetry pipelines that span compute, storage, and networking layers - Familiarity or direct exposure to Kubernetes internals, Slurm, or other distributed orchestrator ecosystems - Experience applying statistical methods, anomaly detection, or time-series forecasting to noisy operational data for predictive maintenance - Knowledge of zero-trust architectures, policy-based access systems, or automated multi-tenant audit streaming

Similar roles