SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Hydrolix is seeking a Senior Site Reliability Engineer to join its Services team and contribute to the reliability and scalability of its cloud data platform built for petabyte-scale datasets. This is a highly technical, hands-on role requiring deep expertise in system reliability, automation, and distributed systems.
Key responsibilities include deploying and maintaining highly reliable Kubernetes clusters and Hydrolix deployments across multiple cloud platforms; designing and implementing systems to enhance service reliability, availability, and performance; building and optimizing CI/CD tools and processes; developing monitoring, alerting, and incident response strategies to minimize downtime; conducting root cause analyses for system failures; and automating repetitive tasks to improve operational efficiency. You will participate in on-call support covering weekday business hours and once-monthly weekend shifts.
You will work closely with software engineering, infrastructure, and product teams to integrate reliability practices throughout the development lifecycle, champion SRE best practices, collaborate with a distributed global team, and interface with customers to resolve incidents.
Required qualifications include a Bachelor's degree in Computer Science, Engineering, or related field; 5+ years of experience as a Site Reliability Engineer or similar role supporting complex distributed systems; experience with monitoring and debugging tools such as Prometheus, Vector, Grafana, Superset, or Kibana; proficiency in at least one major cloud platform (AWS, GCP, Azure, or Linode); experience with SQL databases (PostgreSQL preferred); proficiency in programming languages such as Python, Go, or Rust; strong Linux expertise including performance tuning and system-level troubleshooting; and excellent written and verbal communication skills.