SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ClickHouse, a Forbes Cloud 100 company and leader in real-time analytics and observability, is hiring an experienced Cloud Software Engineer for its Observability Platform team. The role sits at the intersection of distributed systems, cloud infrastructure, and production operations.
You will design, build, and operate distributed systems that ingest, process, and store telemetry at massive scale—processing trillions of events per day with throughput in the hundreds of millions of events per second. The Observability Platform team builds shared systems for telemetry ingestion, durable buffering, processing, storage, autoscaling, and service provisioning. The Internal Observability team operates ClickHouse's company-wide observability platform and partners with engineering teams to improve reliability, debugging, and operational efficiency.
Key responsibilities include:
- Design and operate distributed systems handling telemetry at very high scale
- Own reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage
- Participate in on-call rotation, resolve production incidents, and drive root-cause fixes
- Build software and automation to eliminate repetitive operational work
- Identify architectural bottlenecks and shape the roadmap for next-stage scale
- Collaborate with product, infrastructure, and service teams across ClickHouse
- Contribute to architecture reviews and raise engineering quality
Required experience: 5+ years building and operating production systems at scale; strong Go proficiency; Kubernetes experience; infrastructure-as-code (Terraform, Helm, Argo CD); production experience with AWS, GCP, or Azure; hands-on telemetry systems (OpenTelemetry, Prometheus, Grafana, or similar).
Bonus: ClickHouse experience, high-throughput ingestion/streaming/storage systems, multi-tenant cloud services, infrastructure cost optimization, TypeScript.
You take ownership, debug unfamiliar distributed systems in production, make pragmatic tradeoffs, communicate clearly in async-remote environments, prefer incremental delivery validated in production, and address root causes rather than symptoms.