SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Thought Machine - London, United Kingdom - In-office - posted 2026-09-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Thought Machine is building modern cloud-native core banking and payments technology to replace legacy systems at the world's leading financial institutions. The company has raised over £500m from top-tier investors including JPMorgan Chase, Standard Chartered, and Temasek, and operates globally across London, New York, Singapore, Sydney, and Lisbon with 550+ employees. As a Senior Site Reliability Engineer, you will be a guardian of mission-critical infrastructure powering cutting-edge banking platforms for the world's most influential financial institutions. Based at Thought Machine's London headquarters, you'll join an elite, globally distributed SRE team responsible for running and maintaining robust production infrastructure that serves as the backbone of their SaaS offerings. Key responsibilities include: - Supporting product engineering teams in building fault-tolerant, scalable applications through design discussions, RFCs, and code reviews - Executing department strategies around disaster recovery, backup, redundancy, and capacity planning - Participating in a global on-call rotation to identify and fix bottlenecks in customer environments - Maintaining production systems hosting Vault products - Designing and implementing features that enhance reliability and user experience - Implementing and regularly testing disaster recovery strategies to ensure platform resilience - Maintaining high-quality documentation of assets, processes, and runbooks - Mentoring team members in technical skills and product knowledge You'll work with senior stakeholders across the organization and directly with customers on programs critical to company success. The role involves tackling complex challenges in fleet management automation, designing operational processes that interface between Thought Machine and SaaS customers, and driving the evolution of products through a reliability lens. REQUIREMENTS: - Track record of delivering high-impact projects with focus on long-term scalability, ensuring human intervention scales sub-linearly with usage growth - Up-to-date understanding of design patterns relevant to hosting and networking architectures - Proactive approach to championing product development with desire to build exceptional products, not just solve immediate challenges - High-agency individual who can independently drive projects to completion and effectively delegate work to team members - Strong background in Python or Golang, with experience executing significantly sized projects or initiatives in one of these languages - Experience working with Kubernetes or other container orchestration systems - Experience with automation/configuration management tools (e.g., Terraform, Puppet, Chef, Ansible) - Expertise in one or more of: Database Administration, Networking, Observability Tools (Prometheus, Jaeger), or automation infrastructure - Extensive experience with either GCP or AWS

Similar roles