SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Kraken, a leading crypto exchange platform founded in 2011, is seeking a Site Reliability Engineer to join its Telemetry team. This role focuses on operating and improving the shared telemetry platform that enables teams across the organization to understand, operate, and improve production services at scale.
You will be responsible for maintaining and scaling the infrastructure supporting metrics, logs, traces, alerting, dashboards, and profiling systems. Key responsibilities include operating Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools; managing log pipelines using Vector, Splunk, and Loki; deploying and managing telemetry services using Terraform and container orchestration; and troubleshooting production issues such as missing data, slow queries, broken alerts, and capacity problems.
The role is hands-on and collaborative, requiring you to work with Software Engineers, Platform teams, Security Engineers, and other SREs to improve operational practices. You will participate in incident response, maintain on-call rotations, write runbooks, and build reusable automation and configuration that helps teams manage dashboards, alerts, and telemetry integrations safely.
Required qualifications include 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, or similar production engineering role. You should have demonstrated experience managing production systems at scale that collect, process, store, and serve telemetry; proficiency with Prometheus or Prometheus-compatible monitoring stacks; experience troubleshooting distributed production systems; solid Infrastructure as Code skills (Terraform); experience with containerized workloads (Nomad, Kubernetes); and strong scripting/programming ability. Experience with AI tools and agents to accelerate delivery is valued.
Nice-to-have skills include experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, PromQL, LogQL, Kubernetes operators, Consul, Vault, AWS, and high-volume data pipelines. Background in regulated or financial services environments is a plus.