SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Tabby is a fintech platform that enables flexible payments and financial control for over 25 million users globally. The company processes $18 billion in annual transaction volume and is valued at $6.5 billion, having raised over $1 billion in funding since its 2019 launch.
We are seeking a Senior Database Administrator to join our Infrastructure team and own the reliability and performance of our production ClickHouse clusters. You will work on high-impact projects in a high-growth environment alongside a world-class remote engineering team spread across 20+ countries.
Key Responsibilities:
- Own reliability and performance of production ClickHouse clusters, whether self-hosted on Kubernetes or running on managed services
- Design standard architecture and tooling to scale to dozens of multi-node clusters without proportional operational overhead
- Manage the complete cluster lifecycle as code: provisioning, configuration, scaling, upgrades, and migrations
- Build monitoring and alerting systems that catch problems before users are impacted
- Own backups, disaster recovery, and regularly test restore procedures
- Ensure clusters are secure by default and maintain database security best practices
- Plan capacity and control infrastructure costs
- Serve as the go-to expert on ClickHouse for developers, DevOps engineers, other DBAs, and leadership—from schema design and data ingestion to query performance optimization
- Handle routine database requests and automate repetitive tasks
- Respond to production emergencies outside business hours and lead root-cause analysis and postmortem contributions
- Maintain documentation, runbooks, and share knowledge with the team
Requirements:
- Deep, DBA-level production knowledge of ClickHouse, including understanding how it stores, merges, replicates, and queries data; strong SQL skills including complex analytical query optimization and schema design review
- Ability to design architecture and automation for dozens of multi-node clusters; experience running a large fleet of databases (dozens or hundreds of instances) is a strong advantage
- Production experience running and maintaining stateful workloads on Kubernetes well beyond initial deployment
- Production use of a Kubernetes operator to run databases (CloudNativePG, Zalando Postgres Operator, Altinity, official ClickHouse operator, or equivalent)
- Experience with version upgrades, backup and restore, and migrating large datasets between clusters or regions with zero or minimal downtime
- Strong troubleshooting skills across the full stack: queries, schema design, Kubernetes, Linux, storage, and networking
- Proactive approach to database health; experience building monitoring and alerting with modern observability stacks
- Practical database security knowledge: encryption in transit and at rest, access control, secrets management
- Infrastructure as code with mainstream tools plus at least one programming or scripting language for automation
- Hands-on experience with cloud infrastructure on major providers (not just platform-level usage)
- Clear written and spoken communication with engineers and leadership on architecture, performance, incidents, risks, and costs; ability to independently take problems from investigation to production fix
Nice to Have:
- Hands-on work on Kubernetes database operator source code as a developer or contributor
- Practical use of AI tools for database observability or automating routine operational tasks
- Background in highly regulated environments
- Kubernetes or cloud certifications