SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Braze is seeking a Senior Database Administrator to join the Platform Engineering division within Infrastructure. You will be responsible for operating, scaling, and securing Braze's internal MongoDB platform, which supports product-critical features, internal systems, and distributed services across a growing fleet of MongoDB clusters.
You will partner closely with Site Reliability Engineers (SREs), application engineers, and the platform security team to ensure MongoDB clusters are performant, resilient, compliant, and well-automated. The clusters run on modern Kubernetes-based infrastructure and are managed through GitOps and infrastructure-as-code practices.
Key Responsibilities:
**Cluster Operations & Reliability:** Operate production MongoDB clusters running in Kubernetes across multiple regions and clouds. Monitor cluster health, performance, replication, and storage usage, proactively addressing degradation. Debug and resolve availability, latency, or data consistency issues in partnership with SREs. Participate in on-call rotations as a subject matter expert for MongoDB infrastructure.
**Automation & Tooling:** Automate routine operations such as database user management, cluster scaling, version upgrades, and backup restores. Maintain and improve internal tooling used for MongoDB observability, lifecycle management, and access control. Collaborate with Platform Engineering on Kubernetes-based deployment patterns and GitOps workflows.
**Backup, Restore & Disaster Recovery:** Implement and maintain point-in-time recovery, daily snapshots, and cross-region backup strategies. Regularly validate restores and participate in resilience exercises and disaster recovery planning. Partner with Security and Compliance teams to ensure backups meet RTO/RPO targets and audit requirements.
**Engineering Support & Best Practices:** Assist product engineering teams with schema design, indexing strategies, and performance tuning. Promote internal best practices for efficient queries, data modeling, and shard key design. Document operational runbooks, configuration standards, and performance baselines.
**Requirements:**
- Proven work experience managing MongoDB in production environments at scale
- Strong working knowledge of replica sets, sharding, WiredTiger storage engine, oplogs, and journaling
- Experience operating MongoDB in cloud or containerized environments (Kubernetes preferred)
- Comfortable with automation and scripting in languages like Bash, Python, or Go
- Proficient with MongoDB CLI tools, Atlas APIs, and backup/restore tooling (mongodump, mongorestore, mongosh, Ops Manager)
- Familiar with observability stacks (e.g., Datadog, Prometheus, Grafana) and alerting strategies
- Detail-oriented and reliability-focused; treats data loss and downtime as unacceptable
- Strong communicator able to document clearly, partner cross-functionally, and lead incident response
- Growth-oriented; comfortable evolving systems, adopting new automation, and supporting other engineers
**Bonus Qualifications:**
- Experience with other database platforms (PostgreSQL, MySQL, DynamoDB)
- Familiarity with security best practices and compliance frameworks like SOX or SOC 2
- Experience with GitOps workflows and infrastructure-as-code tools (Terraform, Helm, ArgoCD)
- Contributions to open source MongoDB tooling or community forums