SlipstreamJobsFresh Startup & VC-Backed Jobs

DevOps Engineer

Encord - London, United Kingdom - In-office - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Encord is the universal data layer for AI, helping 300+ AI teams train and run models on the right data. The platform indexes, curates, annotates, and evaluates data across the full AI lifecycle. Trusted by Woven by Toyota, AXA, UiPath, Zipline, and others, Encord has raised $60M in Series C funding and operates as a team of 100+ at the frontier of AI. You will join the platform engineering team in London, embedded in teams building and operating Encord's core infrastructure. Your focus will be ensuring the platform is performant, reliable, observable, and scalable. You'll drive a culture of automation, performance, and resilience through individual contributions and collaboration across multiple squads. Key responsibilities include: - CI/CD & Deployment: Own and continuously improve deployment pipelines; partner with developers to review infrastructure changes, streamline release processes, and champion DevOps best practices across engineering. - Infrastructure & Cloud: Design, deploy, and maintain cloud infrastructure on GCP and AWS; manage Kubernetes clusters, networking, and storage at petabyte scale using infrastructure-as-code. - Automation & Tooling: Drive developer productivity by building, guiding, and reviewing automation and internal tooling; eliminate manual toil. - Performance & Capacity: Profile and optimize services handling large-scale data pipelines; perform capacity planning for storage and compute-intensive workloads; establish performance benchmarks with squads. - Reliability & Availability: Define and own SLIs/SLOs/SLAs for critical services; build alerting, runbooks, and incident response processes; lead postmortems with a blameless culture. - Observability: Instrument services with distributed tracing, logging, and metrics (Prometheus, Grafana, OpenTelemetry, GCP Dashboards); build infrastructure and ensure every service is observable before production. The tech stack includes Python (backend), TypeScript and React (frontend), Kubernetes (deployment), GCP (infrastructure), and PyTorch, CUDA, Ray (machine learning). The company is technology-agnostic and values learning ability. The team works from the London office 4+ days per week, reflecting a strong in-person culture. REQUIREMENTS: - 4+ years of hands-on DevOps, platform engineering, or SRE experience in a production environment - Strong experience building and maintaining CI/CD pipelines and deployment automation at scale - Proven experience with infrastructure-as-code tools (e.g., Terraform, Pulumi) and configuration management - Strong fundamentals in designing, building, and maintaining resilient distributed and/or high-performance systems - Hands-on experience with Kubernetes and containerized workloads in cloud environments (GCP and/or AWS) - Solid understanding of networking, operating systems, and database technologies - Experience with observability fundamentals: metrics, logs, traces, and alerting

Similar roles