SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Webroot is seeking an experienced Manager, Site Reliability Engineering to lead and mature cloud operations on AWS. This is a hands-on leadership role responsible for the stability, security, performance, and continuous improvement of production and non-production environments.
You will lead a Cloud/Production Services function ensuring applications and infrastructure are deployed, operated, and optimized in alignment with architectural standards, security controls, and operational best practices. Working closely with software engineering, security, infrastructure, and customer support teams, you will drive operational excellence while meeting compliance and availability requirements.
Key responsibilities include:
- Lead day-to-day cloud operations for AWS environments, ensuring compliance with security, availability, and governance requirements
- Own the health, performance, and capacity of production and non-production environments
- Provide hands-on technical leadership across cloud infrastructure, applications, and monitoring
- Ensure application deployments and operational practices align with overall architecture and security standards
- Oversee maintenance activities including upgrades, patching, backups, and recovery testing
- Drive proactive monitoring, alerting, and incident management to maintain high availability and reliability
- Act as an escalation point for complex technical issues and production incidents
- Manage and participate in an on-call rotation supporting 24x7x365 operations
- Identify operational inefficiencies and continuously improve processes, tooling, and automation
- Collaborate closely with software development, security, customer support, and infrastructure teams
- Provide technical guidance and mentoring to operational staff
You excel at leading by example in hands-on cloud operations, are self-reliant and proactive, communicate clearly across geographic regions, and are driven by technical excellence. You learn new technologies quickly and can design and maintain effective monitoring and alerting for cloud services.
Requirements:
- 5+ years' experience in application management, systems administration, or cloud operations
- Tertiary qualification in Computer Science or a related discipline (preferred)
- Proven experience managing public cloud environments with strong expertise in AWS
- Familiarity with the Fortify on Demand stack (highly desirable)
- Solid understanding of software architecture and DevOps principles
- Experience with infrastructure automation and scripting (e.g., PowerShell, Octopus Deploy, Terraform)
- Hands-on experience with monitoring and observability tools (e.g., AWS CloudWatch, Nagios, Grafana)
- Strong understanding of cloud security principles, backup strategies, and disaster recovery
- Working knowledge of networking concepts and protocols