SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer (SRE & Platform Reliability)

Affirm - Remote - Remote - posted 2026-08-07

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: PLN 308,000 - 428,000 / annual

Affirm is reinventing credit to make it more honest and friendly, offering consumers flexible buy-now-pay-later options without hidden fees or compounding interest. The Site Reliability Engineering team is a small but critical function that enables engineering partners to "Operate What They Own" with excellence, protecting customer experience. The SRE team defines frameworks and best practices for operating applications, builds tooling, and provides training and consulting across the organization. Key responsibilities include: - Providing data and visibility to teams and leadership on application performance - Guiding the development of SLOs (Service Level Objectives) - Driving the Incident Management and Analysis process - Steering the implementation of Change Management and Deployment practices - Engaging in service and architectural conversations - Recommending observability and alerting configurations In this Senior SRE role, you will own and deliver quarterly goals for your team, lead engineers through ambiguity to solve open-ended problems, and ensure team support throughout delivery. You'll collaborate with infrastructure, product management, and developer experience teams on ideation, technical constraints, and risk-aware decision-making. You'll proactively identify technical solutions and operational processes that strengthen incident readiness, response, and post-incident analysis. You'll support operations and availability of team artifacts through metrics creation and monitoring, and foster a culture of quality and ownership by setting code review and design standards. Additionally, you'll develop talent on your team through feedback, guidance, and leading by example. Required experience includes 4+ years designing, developing, and launching backend systems at scale using languages like Bash, Python, or Kotlin; a track record developing highly available distributed systems using AWS, MySQL, and Kubernetes; meaningful experience contributing to or driving Incident Lifecycle processes; and 4+ years in a Site Reliability or Production Engineering team. You should demonstrate curiosity with empathy, strong opinions loosely held, and experience defining technical plans for significant features with elegant, simple, and extensible designs. Strong verbal and written communication skills supporting global collaboration are essential.

Similar roles