SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Wave is building financial services for Africa, with millions of users across 9 countries. As an Observability Engineer, you'll own and evolve Wave's observability platform across Python backends, GraphQL APIs, Postgres and CockroachDB databases, Kubernetes workloads, cloud infrastructure, and on-premises environments.
You'll report to the Director of Platform and work within the newly formed Performance & Observability team, partnering closely with infrastructure, database, and product engineering teams.
Key responsibilities include:
- Improving production visibility across application code, APIs, databases, caches, Kubernetes, cloud infrastructure, async jobs, and on-premises systems
- Defining meaningful service-level indicators with product and platform teams, reducing alert fatigue, and ensuring alerts are actionable and tied to user impact
- Building internal tooling and self-service workflows to help engineers instrument services, investigate incidents, analyze performance, and understand dependencies
- Identifying reliability, latency, capacity, and cost issues before they become user-facing incidents
- Establishing observability standards, documentation, and training materials across the engineering organization
- Operating and improving the observability platform (Datadog, Honeycomb, Prometheus, Grafana, OpenTelemetry) while controlling costs
You'll need 5+ years in observability, SRE, platform engineering, infrastructure, backend engineering, or production systems. Deep expertise in metrics, logging, tracing, profiling, alerting, dashboards, SLIs, and incident response is essential. You should have experience building internal tools and platforms used by other engineers, excellent communication skills, and pragmatic judgment about tooling trade-offs.
Technical requirements: proficiency in at least one backend language (Python preferred), hands-on experience with observability tools (Prometheus, Grafana, Datadog, OpenTelemetry, Jaeger, Tempo, Loki, Honeycomb, Sentry), familiarity with Postgres, CockroachDB, Redis, GraphQL, or Kubernetes, and experience with OpenTelemetry instrumentation and collector configuration at scale.