SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: EUR 86,000 - 122,000 / annual
Affirm is reinventing credit to make it more honest and friendly, offering consumers flexible buy-now-pay-later options without hidden fees or compounding interest.
The Site Reliability Engineering team is a small but critical function that enables engineering partners to "Operate What They Own" with excellence, protecting customer experience. The team defines frameworks and best practices for operating applications, builds tooling, and provides training and consulting across the organization.
Key SRE responsibilities include providing data and visibility on application performance, guiding SLO development, driving incident management and analysis processes, steering change management and deployment practices, engaging in service and architectural conversations, and recommending observability and alerting configurations.
In this Senior SRE role managing a team, you will own and deliver quarterly goals, lead engineers through ambiguous problems, and ensure team support throughout delivery. You'll collaborate with infrastructure, product management, and developer experience teams on ideation, technical constraints, and risk-aware decisions. You'll proactively identify technical solutions and operational processes that strengthen incident readiness and post-incident analysis.
You'll support operations and availability of team artifacts through metrics creation and monitoring, escalation when needed, and on-call participation. You'll foster a culture of quality and ownership by setting or improving code review and design standards, advocating for them through writing and tech talks, and developing talent on your team through feedback, guidance, and leading by example.
Required experience includes 4+ years designing, developing, and launching backend systems at scale using languages like Bash, Python, or Kotlin; a track record developing highly available distributed systems using AWS, MySQL, and Kubernetes; meaningful experience contributing to or driving incident lifecycle processes; 4+ years in a Site Reliability or Production Engineering team; strong technical planning and code quality skills; experience making impactful changes in large codebases; demonstrated ownership of personal growth; and strong verbal and written communication skills for global collaboration.