SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Modulr is a B2B embedded-payments platform operating at scale in a regulated environment, processing over 200 million transactions and £180 billion in annual payment value. The company is backed by PayPal, FIS, and General Atlantic, with over 400 employees across London, Edinburgh, Amsterdam, Mumbai, and Pune.
As Modulr evolves its core data architecture, this is the first dedicated Database Reliability Engineer role, focused on raising operational resilience of the payments platform. This is a hands-on position for someone with deep database expertise and an SRE mindset—someone who wants to make a large, busy production estate observable, diagnosable, and safe to operate, and to be the database voice shaping how the platform scales. The role carries a clear path toward a data-architecture mandate as the platform matures.
Key responsibilities include:
- Own the reliability, performance, and observability of production databases (primarily PostgreSQL / Amazon Aurora), ensuring issues are identified and understood in minutes, not hours.
- Lead deep-dive diagnosis of production incidents—query performance, lock and wait-event analysis, connection-pool behavior, blocking sessions—and turn findings into lasting fixes and runbooks.
- Build proactive guardrails: query and statement timeouts, connection-pool standards, slow-query and contention alerting, capacity and saturation monitoring.
- Partner with engineering squads to design schemas, indexes, and access patterns that scale, and enforce safe database practices in their services.
- Contribute the database-reliability perspective to significant platform-evolution work, including separation of read and write paths and decomposition of shared database components.
- Improve database observability tooling and dashboards so reliability is measurable and trends are visible before they become incidents.
- Help embed a culture of operational excellence, blameless post-incident review, and continuous improvement.
Requirements:
Essential:
- Substantial hands-on experience operating production PostgreSQL (Aurora / RDS a strong plus) at meaningful scale.
- Genuine SRE mindset: comfortable with observability tooling, metrics and tracing, alerting, automation, and infrastructure-as-code—not just database administration.
- Demonstrable strength in performance diagnosis: reading query plans, identifying lock contention and blocking, resolving connection-pool exhaustion, and tuning for throughput and latency.
- Experience supporting high-availability, low-tolerance systems where correctness and uptime both matter.
- Clear, calm communicator who can lead a live incident and explain root cause to both engineers and non-technical stakeholders.
- Scripting/automation ability (e.g., Python, Bash) and working familiarity with AWS and Kubernetes-based environments.
Desired:
- Exposure to payments, fintech, or other regulated, high-availability domains.
- Experience with data-observability or lineage tooling and with connection-management layers such as PgBouncer / RDS Proxy.
- Familiarity with decomposing monolithic databases and introducing read/write separation.