SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Torq is a Series D-backed AI-native autonomous SecOps platform experiencing explosive growth (200% employee growth, 300% revenue growth). We're seeking a Production Operations Engineer to own production stability end-to-end and drive operational excellence across our engineering organization.
In this role, you'll be the go-to person for production health, proactively identifying risks before they become incidents. You'll own the incident management process, ensuring incidents are properly tracked, follow-ups are driven to completion, and R&D teams write and own their postmortems. Rather than resolving every incident yourself, you'll see what production needs, set direction, and ensure engineering teams follow through.
Key responsibilities include:
- Owning production stability and acting as the primary stakeholder for environment health
- Establishing and managing the incident process, ensuring accountability and continuous improvement
- Partnering with engineering teams to define reliability improvements and ensure prioritization
- Championing production-readiness standards, defining SLOs and error budgets
- Building a culture where teams own the reliability of what they ship
- Communicating clearly with stakeholders across R&D and leadership
You'll need 4+ years in production operations, SRE, DevOps, or related engineering roles. Strong stakeholder management skills are essential—you influence without authority, align teams around priorities, and communicate credibly with both engineers and leadership. You have solid experience with public cloud environments (GCP/AWS/Azure), Docker, Kubernetes, and microservices architectures. Above all, you're driven by ownership, relentless about follow-through, and thrive in collaborative, empowering environments.