SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Trumid is a leading fintech platform revolutionizing fixed-income trading. Founded in 2014, the company operates one of the top three corporate bond e-trading platforms in the U.S., serving 890+ buy-and sell-side institutions with over 1,300 active traders monthly.
This Senior Database Reliability Engineer role owns data resilience and continuity for Trumid's production systems. You will be responsible for ensuring that data-layer failure modes—storage contention, replica lag, region loss—are discovered and mitigated through drills and engineering, not in production incidents. PostgreSQL and RDS form the core of the data infrastructure, and your work ensures this foundation remains uncompromised.
Key responsibilities include:
- Owning durability, recoverability, and performance of the Postgres/RDS fleet across all environments, including replication, failover, backup/restore, and storage behavior under load
- Making recovery provable through restore-tested coverage of every production database, timed failover drills, and measured RPO/RTO per tier
- Managing the complete data lifecycle: retention and cleanup policies, access control, encryption posture, and sensitive data classification
- Hunting performance pathologies at the engine level: lock contention, WAL throughput, replica lag, bloat, index hygiene, and write amplification
- Building database observability and extending AI-assisted operations tooling (health checks, migration PR review)
You will join a team with deep Postgres expertise and work on a production RDS fleet backing a live trading venue. The team has already characterized WAL-write contention under concurrent commits and is actively tuning group-commit behavior, storage classes, and dedicated log volumes based on measurements. Disaster recovery is treated as an engineering discipline with automated cross-region failover measured in minutes and validated through timed, documented drills.
Required qualifications:
- 5+ years running production PostgreSQL at meaningful scale (ideally RDS or Aurora) with depth in internals: replication, WAL mechanics, MVCC, vacuum behavior, query planning
- Full lifecycle data ownership experience: retention, backup/restore, access control, security posture
- Concrete disaster-recovery experience: failovers designed, drills run, measurable outcomes
- Infrastructure fluency: IaC (Terraform), scripting (Python/bash/SQL), Linux, cloud storage/IOPS
- SRE sensibility: SLOs, blameless postmortems, preference for rehearsed over improvised
- Strong technical writing for runbooks and decision records
Nice-to-have skills include warehouse/pipeline experience (BigQuery, AlloyDB, Kafka), regulated/fintech environment exposure, Kubernetes, Prometheus/Grafana/ELK observability stacks, and interest in AI-assisted operations.