SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer, DevOps & Security

Klue - Vancouver, BC, Canada - Hybrid - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Klue is building the future of competitive intelligence with an AI-first platform that generates market, competitor, and buyer insights. The Engineering team is hiring a Senior Software Engineer, DevOps & Security to own the infrastructure and platform that enables the entire engineering organization to build and ship safely at scale. In this role, you will be responsible for making the correct way to build and ship at Klue the easy way. You'll work alongside the DevOps team, engineering managers across Product, Data, and AI, and Security and Compliance stakeholders. You are systems-minded, pragmatic, security-aware, and allergic to toil—someone who jumps in when the platform is on fire but always asks what change stops the next one from starting. Key responsibilities include: - Own the paved road for infrastructure by building and extending the Pulumi component library so teams get correct, disaster-recovery-ready infrastructure by default across GCP environments. - Run GKE like a product, handling cluster and node pool upgrades, in-cluster services, workload right-sizing, and policy enforcement. - Make CI/CD something engineers never think about by owning core building blocks of CI/CD workflows, cutting CI time and flake rate fleet-wide, and standardizing the golden path for shipping services. - Give engineers real visibility into their systems by improving internal developer platform surfaces. - Keep production legible by owning the health of the observability stack: consistent instrumentation, actionable alerts, and reduced on-call noise. - Treat cost like an engineering problem by driving GCP and LLM cost reduction through attribution, right-sizing, commitment planning, and cleanup. - Make the secure path the fast path by extending software supply chain controls, owning infrastructure identity and access, and building SOC 2 and compliance requirements into automated, evidenced controls. Success metrics include: new services provisioned from standardized components without DevOps writing bespoke infrastructure; reactive load and on-call page volume trending down; median pipeline duration and flaky-failure rate dropping across the fleet; controls automated and evidenced with audit findings trending to zero; GCP spend per unit of workload declining; services with published SLOs and real error budgets; and recurring manual work getting automated. Klue is a fast-growing, globally-distributed team based in Vancouver with hubs in Toronto and London. The company is AI-first at its core, with autonomous agents handling market intelligence collection and insight generation. The culture emphasizes curiosity, ownership, and winning together. Requirements: Must-Haves: - Deep, hands-on experience in DevOps, SRE, platform, or infrastructure engineering, operating production systems that paying customers depend on, including being on-call. - Deep, practical Kubernetes knowledge. - Infrastructure as code treated as software (Pulumi or Terraform). - Strong GCP or AWS fundamentals. - Background in systems engineering and writing software. - CI/CD owned as a product; experience building and maintaining pipelines used by other teams (GitHub Actions or Buildkite preferred; GitLab CI / CircleCI transfers). - Hands-on experience with security tooling as an engineering practice. - Compliance engineering for SOC 2 or enterprise partner security programs implemented as automated controls; threat modeling experience. - Ownership under ambiguity; ability to scope vague problems, ship incrementally, and communicate clearly what changed and why. - Fluent with AI coding tools and experience working with AI-first teams. Nice-to-Haves: - Experience owning platform migrations and decommissions. - Built an internal developer platform or self-service tooling that other engineers adopted. - Comfort with LLM infrastructure (gateways, token cost attribution, rate limiting, provider failover). Tech stack: GCP, Kubernetes, Pulumi IaC, Postgres, Temporal.

Similar roles