SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
MX is a fintech company building technology that powers financial applications for banks, credit unions, and fintechs, processing billions of transactions for major financial institutions. The company is in a phase of renewed momentum and scale, with a culture that values curiosity, accountability, and impact.
As a Senior Site Reliability Engineer (titled "Sr. Site Reliability Engineer"), you will build and operate an observability control plane that serves as a multiplier for the entire engineering organization. Rather than building dashboards for each team individually, you establish standards, automate baselines, and shepherd teams toward observability maturity.
Key responsibilities include:
- Build and operate an observability control plane using the Datadog API and Terraform to automate baseline monitors, dashboards, and tagging standards across services.
- Define observability standards for Ruby, Go, and Java services (golden signals, alert quality, dashboard contracts), then audit services against those standards.
- Validate service readiness without owning the alerts and dashboards themselves; service owners maintain their own, and you confirm completeness and correctness.
- Produce monthly observability and service-catalog health reports tracking departed owners, stale dashboards, SLO gaps, and coverage trends.
- Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track progress over time.
- Tune alerting to eliminate false SEV1/2 pages and ensure SEV3/4 alerts are actionable; coach teams on Datadog cost and cardinality.
- Build self-serve onboarding so new services get baseline observability on day one without multi-week embeds.
- Share the team pager: rotate on the shared IR & Observability on-call, triage and investigate live incidents, and take Incident Commander or supporting technical roles as needed.
- Close the detection loop after incidents (gap packs, new monitors, dashboards) so the pager gets quieter over time.
- Run high-value launch and production-readiness reviews as a checkpoint.
The engineering stack includes hybrid infrastructure (AWS and bare metal), services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, data on PostgreSQL and Redis, Datadog for observability, and incident.io for incident response.
Required qualifications: BS in Computer Science or equivalent; 5+ years running production observability, SRE, or DevOps; 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform with Kubernetes proficiency; distributed-systems debugging across microservices; Incident Commander capability; AI- and workflow literacy for scaling reviews and audits.