SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: CAD 93,320 - 138,480 / annual
OpenText is seeking a Senior Site Reliability Engineer to join a globally distributed SRE team responsible for the reliability, performance, and stability of data services powering customer-facing SaaS products. This is a hands-on role focused on operating and scaling distributed data platforms across on-premises and public cloud environments.
You will work with technologies including Kafka, Elasticsearch, Cassandra, Solr, Redis, and OpenSearch across AWS, Azure, and GCP. Key responsibilities include:
- Operating, maintaining, and scaling distributed data services
- Building and enhancing infrastructure across on-premises and public cloud environments
- Developing and maintaining Infrastructure-as-Code using Terraform and Ansible
- Applying patches, performing routine maintenance, and ensuring security and compliance
- Participating in design, deployment, and monitoring of data platforms in collaboration with SRE and engineering teams
- Supporting incident response and participating in on-call rotation for critical services
- Assisting with capacity planning, performance tuning, and health assessments
- Creating and maintaining operational documentation, procedures, and incident reports
- Contributing to automation and reliability initiatives to improve service performance
- Supporting service requests and ensuring SLA/OLA commitments are met
- Participating in team knowledge-sharing and training activities
The role may require shift work and participation in a 24x7 on-call rotation.
REQUIREMENTS:
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field (or equivalent practical experience)
- 4+ years of experience in Information Technology supporting large-scale enterprise systems
- 2+ years operating or supporting distributed data platforms (Kafka, Elasticsearch, Cassandra, Solr, Redis, OpenSearch)
- 2+ years working with automation and configuration tools such as Terraform and Ansible
- Strong knowledge of Linux systems administration
- Experience working with public cloud infrastructure (AWS, Azure, or GCP)
- Solid troubleshooting skills and ability to resolve complex technical problems
- Excellent written and verbal communication skills
- Self-driven, detail-oriented, and able to manage multiple tasks in a fast-moving environment
- ITIL process familiarity (certification is a plus)
- Experience with observability tools (Prometheus, Zabbix, Grafana, New Relic, etc.) is a plus