SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer (Site Reliability Engineer)

Anyscale - Bengaluru, KA, India - In-office - posted 2026-08-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Anyscale is building the commercial platform around Ray, an open-source distributed computing framework used by companies like OpenAI, Uber, Spotify, and Instacart to scale machine learning workloads. The company has raised $250+ million from top-tier investors including Andreessen Horowitz and NEA. As a Site Reliability Engineer, you will ensure the smooth operation of all user-facing services and production systems at Anyscale. You'll work at the intersection of infrastructure, cost optimization, and operational excellence as the company scales. Key responsibilities include: - Develop a unified perspective on cloud component utilization across the company, balancing diverse team needs and identifying cost optimization opportunities. - Ensure deployment methodologies align with reliability goals and implement best practices for production environments. - Build robust observability infrastructure for metrics, logging, and tracing to enable quick issue identification and resolution. - Create monitoring and alerting systems at multiple levels, allowing teams to contribute to and enhance monitoring capabilities. - Establish testing infrastructure to support effective test writing and execution across teams. - Develop tools for measuring and defining service level objectives (SLOs) at the organization level. - Implement on-call systems and incident management processes, continuously improving incident response capabilities. - Coordinate cloud service creation and deployment, tracking deployments and establishing effective communication channels. You'll apply sound engineering principles, operational discipline, and mature automation to Anyscale's environments and codebase. This role requires at least 3 years of relevant experience in site reliability engineering or a similar infrastructure/operations role.

About Anyscale

AI / Data / Infrastructure — distributed computing and AI workload platform built around Ray.

Similar roles