SlipstreamJobsFresh Startup & VC-Backed Jobs

Engineering Manager, Infrastructure Engineering

CoreWeave - New York, NY, United States - Hybrid - posted 2026-07-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

CoreWeave is seeking an experienced Engineering Manager to lead the MetalDev RAS (Reliability, Availability & Serviceability) team within Hardware Engineering Dev. This is a first-line management role where you will build, coach, and grow a team of infrastructure and site reliability engineers responsible for the reliability and performance of CoreWeave's bare-metal infrastructure at scale. You will own both team health and system health—setting technical direction, driving operational excellence, and partnering with cross-functional teams and external vendors to deliver highly performant and resilient infrastructure solutions. The MetalDev RAS team is responsible for the software and infrastructure that keeps CoreWeave's fleet of GPU servers reliable, available, and serviceable at massive scale, sitting at the intersection of software engineering, infrastructure, and hardware. Key responsibilities include hiring and developing infrastructure/SRE engineers; establishing and improving incident response processes, runbooks, and RCA/PIR practices; defining and driving KPIs, SLAs, and SLOs; championing system observability using tools like Prometheus and Grafana; setting automation strategy to reduce manual toil; and leading communication during major incidents. You will collaborate across engineering teams on platform reliability, resilience improvements, and disaster recovery, while representing the team in planning and prioritization. Required qualifications: 3+ years of engineering management experience leading software/infrastructure/SRE teams (first-line management) with a strong individual-contributor background; 7+ years combined experience in cloud operations, SRE, infrastructure, or related technical roles; strong understanding of cloud platforms, Kubernetes, and bare-metal infrastructure; solid grounding in incident management practices (ITIL, SRE best practices); and experience leading teams developing software in Go or comparable systems languages.

Similar roles