SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 199,750 - 270,000 / annual
Horizon3.ai is a fast-growing, remote-first cybersecurity company focused on autonomous penetration testing and vulnerability assessment. The NodeZero platform enables organizations to proactively discover and verify exploitable attack vectors at scale across internal, external, cloud, and hybrid environments. The company serves organizations from educational institutions to Global 100 enterprises, including IT Ops/SecOps teams, consulting pentesters, MSSPs, and MSPs.
You will own and evolve the engineering-wide SRE strategy, operating model, and reliability standards, aligning them to customer impact, business priorities, and risk. This is a foundational, hands-on role for an experienced engineer who will set technical direction across teams, lead high-impact reliability initiatives, and establish practices and systems enabling safe production operations.
Key responsibilities include:
- Design and evolve engineering-wide SRE strategy, operating model, and reliability standards
- Lead cross-functional alignment across Infrastructure, product, service, security, and business teams to improve reliability, observability, incident response, and operational readiness
- Establish organization-wide service ownership, SLIs, SLOs, and error budgets for critical customer paths
- Define and drive adoption of observability standards across pipelines and platform components
- Set standards for dashboards, actionable alerting, runbooks, and escalation paths
- Drive end-to-end complex cross-functional reliability initiatives
- Raise the engineering-wide bar for incident management, incident command, on-call health, post-incident learning, and recovery readiness
- Shape the technical direction, operating model, and growth path of the SRE function
- Participate in 24/7 on-call rotation and help design a sustainable on-call model
The role is fully remote with up to 10% travel for team off-sites and in-person project kick-offs.
REQUIREMENTS:
- Experience designing, operating, and troubleshooting large-scale distributed systems in production environments
- Deep knowledge of reliability engineering, observability, incident management, and production operations, with demonstrated ability to turn that knowledge into standards and practices adopted by others
- Experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership
- Backend experience building backend systems and automation that reduce operational toil, strengthen safeguards, and improve operational efficiency
- Experience leading high-severity incidents and improving incident response programs
- Excellent written and verbal communication skills including technical designs, runbooks, postmortems, and operational documentation
- Python and Terraform (Infrastructure as Code) or equivalent automation and infrastructure-as-code tools
- Experience with observability tools such as Datadog, New Relic, Grafana, or equivalent platforms
- Experience operating production services in AWS and Kubernetes
- Experience with CI/CD pipelines such as Gitlab CI, ArgoCD, or GitOps workflows