SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
CoreWeave, a publicly traded AI infrastructure company (Nasdaq: CRWV), is seeking a Senior Infrastructure Engineer to join the Hardware Engineering Dev team. Reporting to the Engineering Manager, you will own the development, deployment, and monitoring of services that manage bare-metal infrastructure at scale.
Key responsibilities include leading incident response efforts and coaching junior team members through resolution. You'll conduct root cause analysis, drive post-incident reviews, and develop incident response playbooks to prevent service degradation. You'll own system observability using Prometheus and Grafana to proactively detect performance bottlenecks and lead automation efforts to minimize manual intervention.
You'll build strategy around making core services perform optimally at scale, defining KPIs and SLAs for incident management aligned with organizational reliability objectives. Collaboration is central: you'll work across teams to improve platform reliability, resilience, and disaster recovery. You'll design and implement solutions for operational efficiency, create CI/CD pipelines, ensure smooth server hardware lifecycle management, and partner with Fleet Operations to design scalable tooling that enables self-service and reduces escalation overhead.
Additional responsibilities include building dashboards and alerts for efficient troubleshooting, participating in on-call rotation, triaging support channel issues, and driving reduction in on-call queries and incidents over time.
Required qualifications: 7+ years in cloud operations, SRE, or related technical roles. Strong understanding of cloud platforms (Kubernetes, AWS, GCP) and cloud infrastructure. Familiarity with incident management practices (ITIL, SRE best practices). Proficiency in Go or Python. Prior experience with Prometheus/Grafana and Kubernetes. Excellent documentation skills, strong analytical and problem-solving abilities, and production on-call rotation experience.