SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Obsidian Security, a leading SaaS security platform trusted by 200+ global enterprises including Snowflake and T-Mobile, is seeking a Site Reliability Engineering Lead to establish and lead production reliability capabilities within its growing Taiwan engineering organization.
You will be responsible for the reliability, security, scalability, and operational effectiveness of cloud services protecting some of the world's largest enterprises. This is a hands-on leadership role where you will lead work across service reliability, cloud infrastructure, observability, incident management, capacity planning, vulnerability remediation, and production readiness.
Key responsibilities include:
- Establish and lead the SRE function in Taiwan, including technical roadmap, operating model, hiring plan, and global team relationships
- Improve availability, performance, scalability, security, and cost efficiency of production services
- Define service-level indicators, objectives, error budgets, and operational health metrics
- Build and improve observability across applications, data pipelines, APIs, infrastructure, and customer-facing workflows
- Lead production incident response, technical coordination, and service recovery
- Establish on-call practices, escalation paths, runbooks, and incident command processes
- Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes
- Develop automation to reduce manual operations, deployment risk, and repetitive work
- Partner with engineering teams on production readiness, resilience testing, and safe service rollout
- Improve deployment and change-management practices through progressive delivery and automated validation
- Own or coordinate infrastructure vulnerability and CVE management
- Partner with Security and Engineering to strengthen cloud configuration and infrastructure security
- Identify architectural weaknesses and lead cross-team improvements
- Mentor SREs and software engineers in reliability engineering and production operations
- Collaborate with teams across Taiwan, US, UK, and Australia for global operational coverage
You will work closely with the Director of Engineering – Taiwan and global engineering, infrastructure, security, and product teams. As the Taiwan site grows, you will recruit and develop a high-performing SRE team and help create a strong, shared operational culture.