SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anyscale is commercializing Ray, an open-source distributed computing framework used by leading companies like OpenAI, Uber, Spotify, and Instacart to scale machine learning workloads. The Ray Core team develops and maintains the C++ backend, including the distributed scheduler, language runtime integration, I/O and memory subsystems, and is responsible for the reliability, scalability, and performance of the entire platform.
In this role, you will lead cross-team projects while mentoring junior engineers, develop high-quality open-source software to simplify distributed programming, and identify and implement architectural improvements to Ray's core systems. You'll work on a balanced mix of new features, distributed libraries, test infrastructure improvements, debugging, and longer-term architectural enhancements.
Key projects include optimizing performance of large-scale workloads, building stability and stress testing infrastructure, and improving fault tolerance and high availability. You'll also communicate your work to the broader community through talks, tutorials, and blog posts.
Required qualifications include at least 5 years of relevant work experience, proven expertise in building scalable and fault-tolerant distributed systems, extensive C/C++ experience with low-level operating systems knowledge, and solid fundamentals in algorithms, data structures, and system design. Preferred qualifications include knowledge of distributed model training and inference techniques (tensor parallel, pipeline parallel) and GPU programming experience.
Anyscale is backed by Andreessen Horowitz, NEA, and Addition, with $250+ million raised to date.
About Anyscale
AI / Data / Infrastructure — distributed computing and AI workload platform built around Ray.