SlipstreamJobsFresh Startup & VC-Backed Jobs

Ingénieur.e staff de fiabilité des sites

QuoteMachine - Montreal, QC, Canada - Hybrid

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

As a Staff Site Reliability Engineer, you will be the technical backbone of the Data Office's infrastructure platform, responsible for reliability, scalability, and developer experience across the entire Data business unit. You will solve systemic platform-level problems rather than individual tickets, acting as a force multiplier for the engineering team around you. Your responsibilities span four key areas: (1) Hands-on engineering and platform ownership (70%) — design and deliver major infrastructure improvements, write production-quality infrastructure-as-code using Terraform, drive solutions from design through delivery, and tackle the most complex platform challenges including data infrastructure for batch and streaming workloads, ML/AI environments (Vertex AI, model serving, GPU-optimized compute), and BI service layers (Looker infrastructure and GCP integration); (2) Team coordination and technical leadership (10%) — lead architecture and design discussions for cross-functional projects, review critical requests, facilitate solution design sessions, and provide technical guidance that maintains team coherence and momentum; (3) General meetings and ceremonies (10%) — lead agile ceremonies, incident reviews, and cross-functional sync meetings, moving from participation to leadership; (4) Mentoring and technical review (10%) — actively mentor senior and mid-level SRE staff, establish standards for IaC quality, observability practices, and operational discipline through code reviews, pair programming, and documentation. You will also participate in on-call rotation and incident response, and contribute to broader organizational goals as needed. Required expertise includes deep GCP infrastructure knowledge (compute, networking, IAM, GKE, data services, FinOps), mastery of Terraform as primary IaC tool, hands-on experience with Looker infrastructure and ML/AI platform tools (Vertex AI, model serving, training pipelines), strong Bash and Golang skills (Python a plus), solid observability experience (metrics, logs, traces, alerts, SLO/SLI design, DataDog), product-minded culture emphasizing platform usability, CI/CD pipeline experience (GitHub Actions, Circle CI, GCP Cloud Build), ability to communicate infrastructure tradeoffs to both technical and non-technical stakeholders, and AI competency as a go-to reference within the Data business unit. You should evaluate design decisions through a security and compliance lens, demonstrate self-awareness with a commitment to continuous learning, and have strong mentoring and coaching abilities.

Similar roles