SlipstreamJobsFresh Startup & VC-Backed Jobs

Lead Software Engineer - Site Reliability

Freshworks - Hyderabad, Telangana, India - In-office - posted 2026-09-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Freshworks is seeking a Lead Site Reliability Engineer to design and maintain systems that prioritize uptime, performance, and observability at scale. You will partner with engineering, platform, and product teams to shift reliability left and establish performance and availability standards across the organization. Key Responsibilities: - Design and implement tools to improve availability, latency, scalability, and system health - Define SLIs/SLOs, manage error budgets, and drive performance engineering initiatives - Build and maintain automated monitoring, alerting, and remediation pipelines - Collaborate with engineering teams to improve reliability by design - Lead incident response, root cause analysis, and blameless postmortems - Champion observability across services—logs, metrics, and traces - Contribute to infrastructure architecture, automation, and reliability roadmaps - Advocate for SRE best practices across teams and functions Requirements: - 7–12 years of experience in SRE, DevOps, or Production Engineering roles - Strong coding proficiency with ability to develop clear, efficient, well-structured code - In-depth Linux expertise for system administration and advanced troubleshooting - Practical experience with Docker and Kubernetes for application deployment and management - Ability to design, implement, and maintain CI/CD pipelines - Understanding of security best practices and compliance in infrastructure - Experience designing and implementing highly available, scalable, and resilient distributed systems - Proficiency in Infrastructure as Code (IaC) tools and infrastructure automation - Deep knowledge and practical experience with Disaster Recovery (DR) and High Availability (HA) strategies - Hands-on experience implementing monitoring, logging, and tracing tools - Strong system design skills for distributed systems with focus on reliability and operations - Excellent analytical and diagnostic problem-solving skills - Degree in Computer Science, Engineering, or related field - Proven track record building and scaling production services with high uptime targets (99.99%+) - Clear history of reducing incident frequency and improving response metrics (MTTD/MTTR) - Strong communication skills and ability to thrive in high-pressure environments - Passion for automation, chaos engineering, and operational excellence

Similar roles