SlipstreamJobsFresh Startup & VC-Backed Jobs

Database Reliability Engineer - DBRE

Cognite - Bengaluru, India - In-office - posted 2026-08-12

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Cognite is seeking a Senior Database Reliability Engineer to join the Cloud Deployment team and own the reliability, scalability, automation, and operational excellence of core database infrastructure supporting industrial AI and data solutions. You will work across PostgreSQL, Elasticsearch, and Kafka, supporting a multi-cloud, Kubernetes-based platform at significant scale. This role is ideal for someone who enjoys solving complex infrastructure and database challenges, eliminating operational toil, and building highly automated systems. Managed PostgreSQL Fleet Orchestration: Standardize and automate lifecycle management of 1000+ PostgreSQL instances across Azure, AWS, and GCP. Build and maintain infrastructure using Infrastructure as Code (IaC) and automation frameworks. Automate provisioning, configuration, patching, upgrades, backups, and operational workflows. Work with cloud-managed PostgreSQL services including Azure Database for PostgreSQL – Flexible Server, Amazon RDS, and GCP Cloud SQL. Evolve provisioning and configuration using Kubernetes-based and Terraform-driven workflows. Hybrid Elasticsearch Operations: Design, operate, and scale high-performance Elasticsearch clusters across Elastic Cloud (SaaS) and ECK (self-managed Kubernetes) environments. Own reliability and performance of Elasticsearch infrastructure supporting search, semantic retrieval, vector workloads, and operational data. Establish best practices around cluster sizing, capacity planning, shard management, upgrades, backups, monitoring, and disaster recovery. Build automation to simplify provisioning and lifecycle management. Kafka Streaming Infrastructure: Own reliability, scalability, and performance of Kafka clusters supporting high-throughput, low-latency event streaming. Operate Kafka in both self-managed (Kubernetes-based, e.g., Strimzi/Kafka Operator) and managed (e.g., Confluent Cloud, MSK) configurations across multi-cloud environments. Drive capacity planning, partition and topic design, replication strategy, and broker-level performance tuning. Establish and automate best practices around upgrades, rolling restarts, backups/disaster recovery, and schema evolution. You will partner closely with Software Engineering, SRE, Platform Engineering, and Product teams to ensure data platforms are highly available, scalable, secure, and resilient.

Similar roles