SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pliant is a Berlin-based European fintech specializing in B2B payment solutions. The company provides a modular, API-first platform helping businesses streamline spending, improve cash flow, and integrate payments into financial workflows. Pliant serves over 4,000 businesses and 20+ partners globally as a licensed e-money institution (EMI), issuing credit cards in 11 currencies across 30+ countries.
You will be the first hire for a Site Reliability function at Pliant, building this practice from the ground up. Currently, reliability responsibilities are scattered across teams with no unified framework, minimal on-call structure, and no SLO standards. Your mission is to standardize reliability practices while hiring and coaching engineers to embed reliability into the software development lifecycle.
Key responsibilities include: defining SLO and error budget frameworks for product teams; owning blameless post-mortems and root-cause analysis; implementing production readiness reviews; closing observability gaps in Datadog (alerts, dashboards, on-call noise); and building the team from scratch with strong technical and cultural standards.
You bring 7-10 years of engineering experience with at least 3 years directly managing engineers. You have a hands-on production/reliability background (this is not a first management role), strong AWS and Terraform expertise, and experience building on-call rotations and incident management processes. You're skilled at platform observability, comfortable with AI-assisted development tools (Claude, Cursor), and have a track record of pushing reliability practices upstream into product teams.
In the first year: establish on-call rotation and incident process from scratch while hiring initial team members (months 1-3); define SLOs for critical services, implement incident review process, and hire at least one additional engineer (by mid-year); establish Site Reliability as a trusted function across the organization with declining repeat incidents (by year-end).
The stack includes Terraform, Spacelift, AWS (with dedicated PCI-scoped account), and Datadog. The company operates under PCI DSS, SOC 2, and ISO 27001 compliance requirements, with Platform Core migrating toward Kubernetes.