SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

Earnin - Mountain View, CA, United States - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 252,000 - 308,000 / annual

EarnIn is seeking a Staff Site Reliability Engineer to lead the company's next stage of reliability maturity. As a pioneer in earned wage access, EarnIn serves millions of community members who depend on fast, reliable, and trustworthy financial products. This role is critical to scaling reliability practices across the engineering organization without relying on heroics or tribal knowledge. You will serve as a Staff-level technical leader, defining reliability standards and architecture across critical services. The core mission is to embed an AI-first operating model into EarnIn's reliability practices—using AI to actively detect, investigate, respond to, learn from, and prevent production issues while maintaining human ownership and engineering judgment at the center. Key responsibilities include: **Reliability Strategy & Standards**: Define and evolve SLIs, SLOs, error budgets, production readiness, observability, incident response, and resilience patterns. Establish a reliability operating model that clarifies service ownership and operational expectations. Use AI-assisted analysis to interpret trends, detect weak signals, highlight capacity risks, and generate actionable reliability scorecards for teams. **AI-First Incident Response**: Overhaul the incident lifecycle to achieve faster detection, sharper triage, and richer context retrieval. Command high-severity incidents as Incident Commander. Design workflows where AI assists with alert correlation, signal enrichment, root-cause exploration, runbook retrieval, postmortem drafting, and corrective-action tracking—while ensuring all critical steps remain reviewable and auditable with human verification. **On-Call Quality & Toil Reduction**: Elevate on-call by silencing noisy alerts, automating repetitive investigations, and enabling responders to rapidly digest service context. Build tools that gather context from Datadog, CloudWatch, incident.io, Slack, runbooks, deployment history, and service metadata. Transition teams from reactive paging to proactive reliability enhancement. **Architecture & Resilience**: Steer service designs for graceful degradation, failure isolation, robust capacity planning, and operational safety across EarnIn's AWS environment (EKS, Kafka, DynamoDB, RDS, SQS). Apply production data and incident learnings to spot architectural risks before they recur. **Mentorship & Cross-Org Influence**: Coach engineers in reliability practices, incident response, SLOs, observability, production debugging, and AI-assisted workflows. Direct design reviews, incident reviews, and operational maturity discussions. Produce documentation, tooling, and reusable patterns that unlock reliability knowledge. You will work closely with SRE, product engineering, infrastructure, security, and leadership teams to embed reliability practices that scale and make AI-assisted operations the default path for every team that owns a service. **Requirements:** - 7+ years in SRE, Software Engineering, or Infrastructure Engineering with increasing scope and cross-org influence - Demonstrated track record of KPI-driven reliability and operational excellence improvements at scale (MTTR, MTTD, alert quality, incident recurrence, SLO attainment, on-call health, corrective-action completion) - Shipped experience applying AI/LLMs to engineering or operational workflows (alert triage, runbook automation, incident investigation, postmortem drafting, remediation recommendation, operational knowledge retrieval, or agentic operations tooling) - Significant expertise with SLIs, SLOs, error budgets, incident command, blameless postmortems, and recurrence prevention in large-scale distributed systems - Strong software engineering ability in Python, Go, or similar languages; you build tools and automation, not just dashboards - Deep observability experience with Datadog, CloudWatch, OpenTelemetry, or similar platforms, with bias toward signal-heavy alerting - Strong infrastructure-as-code and cloud infrastructure experience (Terraform, Kubernetes, AWS, safe and reversible deployment practices) - Practical experience using AI-assisted development tools (Cursor, Claude Code, Copilot, ChatGPT, or similar) to accelerate your own work and model adoption for teams - Plus: Experience in fintech, regulated environments, SOC 2, PCI, FinOps, or cost/performance tradeoffs in high-scale systems

Similar roles