SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Griffin is a fully regulated UK bank founded in 2024 that powers the accounts and payments offerings of product-led fintechs and platforms including Uber, Yonder, Prosper, and Sidekick. The company provides banking infrastructure for remittance, payroll, wealth, insurance, proptech, payments, neobanking and related services, with plans to expand into card issuing, stablecoin infrastructure, and agentic finance.
The Infrastructure team sits within the Craft function and consists of three people—two engineers and an engineering manager—working closely with the CTO and Founder. The team is responsible for building the foundational systems, tooling, and automation that enable Griffin to run a highly resilient and reliable platform while allowing other teams to ship quickly.
In this role, you will manage the infrastructure the bank runs on, spanning both cloud environments and hardware in third-party data centers. Day-to-day responsibilities include managing FoundationDB clusters and building tooling for hitless upgrades, managing Kubernetes, writing infrastructure as code in Pulumi and CDK, running the Bazel build farm, and owning the deployment pipeline from CI to production. Operational resilience is a core focus: you'll design and implement replica clusters, multi-region failover strategies, and chaos engineering tests to validate system behavior under failure conditions. You'll also manage security access controls, CircleCI pipelines, DataDog integration, and cloud spend optimization.
Current roadmap priorities include building out self-hosted infrastructure to support hybrid-cloud setups and reduce single-vendor reliance, improving Kubernetes scalability, enhancing the build farm, establishing PKI infrastructure, and implementing a global HTTP routing layer for seamless failover between sites.
Success in this role means thinking ahead to identify infrastructure issues before they become production problems. Examples include eliminating entire classes of incidents through architectural fixes, proving the database and compute platform can handle 10x traffic volumes through stress testing, cutting deploy times in half by rethinking the pipeline, validating multi-region failover under realistic scenarios, redesigning service discovery to eliminate recurring networking issues, and ensuring compliance with regulatory requirements like FCA standards through proactive infrastructure design.
You will own infrastructure architecture decisions, production reliability across database, compute, build farm and deployment pipelines, the CI/CD and monitoring systems other teams depend on, and infrastructure cost and efficiency across AWS and DataDog.
The tech stack includes FoundationDB, Kubernetes, Bazel, Pulumi, CDK, CircleCI, DataDog, Clojure, and TypeScript. Griffin is remote-first since its establishment, with asynchronous and fully flexible work. The company has a small London office in Moorgate for those who want it, but there is no expectation to be there. The culture emphasizes high-trust, high-autonomy environments where smart, motivated people thrive without hand-holding.
REQUIREMENTS:
- Solid systems experience: production container orchestration, infrastructure as code, IP networking debugging
- Cloud infrastructure experience: building and managing production environments, networking, identity management, compute at scale
- Programming fundamentals: ability to write code to solve problems. TypeScript, Python, Go, or functional languages preferred. Bash scripting and YAML alone are insufficient.
- Comfort with startup environments: small teams, shifting priorities, high autonomy
- Willingness to learn Clojure (not required on entry, but essential for the role). Prior functional language experience is helpful.
- Ability to pick up new languages and approaches; comfort reading diverse codebases of open-source systems and tools.