SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Astronomer is seeking a Customer Reliability Engineer (Infrastructure) to join the CRE team in Hyderabad. The CRE team is responsible for the success of customers using Astronomer's managed Airflow service, focusing on operating, monitoring, and maintaining the platform to ensure availability, predictability, and reliable operations.
In this infrastructure-focused role, you will specialize in the reliability of underlying cloud infrastructure and Kubernetes clusters. You will respond to incidents raised by customers or detected by monitoring systems, troubleshoot customer environments, and take steps to ensure permanent resolution or ongoing monitoring. As an owner of the observability platform, you'll have unlimited potential to improve platform reliability and deliver exceptional customer outcomes.
This is a directly customer-facing role offering exposure to diverse problems and requirements. You'll interface with customers across various industries, cloud providers, and different expectations. Your contributions will directly impact customer success with Astronomer products and enable meaningful improvements to the customer experience.
Key responsibilities include:
- Providing solutions to customers for successful product usage
- Troubleshooting customer environments and active incident triaging
- Providing feedback to product development teams on customer needs and pain points
- Building and maintaining monitoring and alerting systems
- Creating automation for efficient daily operational task handling
- Contributing to product architecture decisions
- Owning the customer experience by working directly with customers to prioritize issues, meet SLAs, and provide guidance on production readiness
- Participating in a fully distributed team environment
- Enhancing customer documentation
- Working on modern, cloud-native products that integrate with dozens of systems
- Maintaining 24x7 coverage through a specified 6-hour pager period during work hours
- Participating in paid on-call rotation for weekend coverage
Requirements:
- 3+ years of experience with large, complex cloud infrastructures operating at scale
- 2+ years of experience with Kubernetes
- Experience managing production distributed systems with at least one major cloud provider (AWS, GCP, or Azure)
- Strong networking experience with one of the major cloud providers
- Strong Linux experience
- Knowledge of operating and monitoring distributed systems
- Experience with observability tools
- Previous experience handling customer issues (internal and external)
- Strong communication skills
- DevOps or CI/CD experience
- Python scripting
Bonus qualifications:
- Site Reliability Engineer (SRE) experience
- Experience with Kubernetes Custom Resources
- Airflow or Big Data Orchestration experience
- Infrastructure as Code (IaC) experience
About Astronomer
AI / Data / Infrastructure — data orchestration platform and commercial home of Apache Airflow.