SlipstreamJobsFresh Startup & VC-Backed Jobs

Cloud Software Engineer - Observability Platform

ClickHouse - Remote - Remote - posted 2026-07-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 141,000 - 220,000 / annual

ClickHouse, a Forbes Cloud 100 company and leader in real-time analytics and observability, is hiring a Cloud Software Engineer for its Observability Platform team. The role sits at the intersection of distributed systems, cloud infrastructure, and production operations. You will design, build, and operate distributed systems that ingest, process, and store telemetry at massive scale—processing trillions of events per day with throughput in the hundreds of millions of events per second. The Observability Platform team builds shared systems for telemetry ingestion, durable buffering, processing, storage, autoscaling, and service provisioning. The Internal Observability team operates ClickHouse's company-wide observability platform and partners with engineering teams to improve reliability and operational efficiency. Key responsibilities include: - Design and operate distributed systems for high-scale telemetry pipelines - Own reliability, performance, capacity, and cost-efficiency of telemetry infrastructure - Participate in on-call rotation, resolve production incidents, and drive root-cause fixes - Build software and automation to eliminate repetitive operational work - Identify architectural bottlenecks and shape the roadmap for scaling - Collaborate with product, infrastructure, and service teams - Contribute to architecture reviews and raise engineering quality You take ownership of systems, debug unfamiliar distributed systems in production, make pragmatic tradeoffs between reliability and speed, and communicate clearly in a remote async environment. You prefer incremental delivery and address underlying causes of operational problems. Required: 5+ years building and operating production systems at scale; strong Go proficiency; Kubernetes experience; infrastructure-as-code (Terraform, Helm, Argo CD); production cloud experience (AWS, GCP, or Azure); hands-on telemetry systems (OpenTelemetry, Prometheus, Grafana). Bonus: ClickHouse experience, high-throughput ingestion/streaming systems, multi-tenant cloud services, infrastructure cost optimization, TypeScript.

Similar roles