SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 202,500 - 274,000 / annual
Intuit is seeking a Staff App Ops Engineer to lead operational excellence and site reliability for its Fintech platform serving millions of small business customers globally. This is a hands-on technical leadership role within the Fintech Systems Engineering and Operational Excellence team.
Key Responsibilities:
- Own "always-on" reliability at scale, driving operational excellence for distributed systems on AWS Cloud with 5-9's availability targets serving millions of customers
- Architect cloud-native data systems with high availability, security, and performance at multi-million-user scale
- Design and ship AI-powered AIOps tools that auto-detect, triage, and remediate issues to multiply team impact
- Lead incident response and production support as part of on-call rotation, engaging and resolving production issues
- Champion resilience best practices including FMEA validation, chaos engineering, and disaster recovery runbooks
- Elevate operational excellence through best-in-class monitoring and resolution techniques to reduce MTTD/MTTR
- Automate relentlessly to reduce developer toil and manual steps in standard operating procedures
- Drive complex cross-team initiatives, sequencing dependencies and managing risk to deliver ambitious programs end-to-end
- Influence beyond your team by partnering with platform and leadership to remove organizational blockers and accelerate adoption of Intuit-wide standards
- Shape engineering culture through active participation in RCAs and contributing to best practices
Required Qualifications:
- 8+ years of hands-on development and operational experience building and maintaining infrastructure in AWS
- Deep expertise in at least one SRE discipline (automation, monitoring, DevOps, or cloud operations)
- Deep AWS and Kubernetes expertise including cloud hosting, high-availability architectures, DR strategies, Docker, Kubernetes, and ArgoCD
- Proficiency in modern programming/scripting languages (Go, Java, Python, or Ruby) for application development and DevOps automation
- Extensive observability and performance experience with tools like Splunk, Wavefront, AppDynamics, Prometheus, or distributed tracing
- Strong grasp of SSDLC and CI/CD pipelines with judgment to build and evolve them safely at scale
- Proven technical leadership with data-backed evidence of impact and track record of driving initiatives independently
- Bias for action, taking initiative and delivering work incrementally
- Sharp production diagnostic skills and passion for resolving pre-production and production issues under pressure
- Strong collaboration and communication skills