SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ClickHouse, a Forbes Cloud 100 company and leader in real-time analytics and observability, is hiring an experienced Cloud Software Engineer for its Observability Platform team. The role sits at the intersection of distributed systems, cloud infrastructure, and production operations.
You will design, build, and operate distributed systems that ingest, process, and store telemetry at massive scale—the company's platforms process trillions of events per day at hundreds of millions of events per second. Key responsibilities include owning the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems; participating in on-call rotations and resolving production incidents; building software and automation to eliminate repetitive operational work; identifying architectural bottlenecks; and collaborating with product, infrastructure, and service teams across the organization.
The Observability Platform team builds shared systems for telemetry ingestion, durable buffering, processing, storage, autoscaling, and service provisioning. The Internal Observability team operates ClickHouse's company-wide observability platform and works with engineering teams to improve reliability and operational efficiency. You will take ownership from design through operation, debug unfamiliar distributed systems in production, make pragmatic tradeoffs between reliability and cost, and address root causes of operational problems.
Required: 5+ years building and operating production systems at scale; strong Go proficiency; experience with Kubernetes; infrastructure-as-code (Terraform, Helm, Argo CD); production experience with AWS, GCP, or Azure; hands-on experience with telemetry systems (OpenTelemetry, Prometheus, Grafana). Bonus: ClickHouse experience, high-throughput ingestion/streaming systems, multi-tenant cloud services, infrastructure cost optimization, TypeScript.