SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Adyen is a global financial technology platform providing payments, data, and financial products to enterprise customers like Meta, Uber, H&M, and Microsoft. The Observability team builds and maintains products and services that enable engineers across Adyen to understand service behavior in real-time, diagnose issues reliably, and discover data through a unified observability ecosystem. The platform processes billions of telemetry events daily across logs, metrics, and traces.
As a Staff Software Engineer on the Observability team, you will play a key architectural role shaping how teams work with telemetry data and driving strategic decisions. You will define and lead logging, metrics, tracing, and alerting strategies across the platform. Key responsibilities include solving scaling bottlenecks in critical telemetry data pipelines, creating advanced tooling to accelerate root-cause analysis and reduce mean time to resolution, and guiding software engineering and SRE teams on monitoring best practices and code instrumentation.
You will identify and tackle architectural bottlenecks and single points of failure to improve observability platform reliability and user experience. You will drive feature ideation, implementation, and adoption while collaborating on planning and refinement to ensure engineering alignment and reduce complexity. On-call duties include responding to alerts and providing production support to other engineers.
Required qualifications include 5+ years of experience with highly distributed systems, expertise in designing and implementing APIs and data pipelines for high-throughput real-time data ingestion, and experience improving software reliability across availability, performance, latency, efficiency, and capacity. You must be proficient in Go and/or Java, have advanced system design knowledge, strong stakeholder management and mentorship capability, and hands-on experience operating core telemetry data stores at scale (Elasticsearch, Opensearch, VictoriaLogs, Clickhouse, Prometheus, VictoriaMetrics, Grafana Tempo, OpenTelemetry, Alertmanager). Experience with containerization (Docker, Kubernetes) and infrastructure-as-code tools (Terraform) is required.
Nice-to-have qualifications include experience with highly available fault-tolerant replicated data storage systems, large-scale data processing systems, infrastructure and platform experience, and contributions to open-source observability projects.