SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Buildkite is hiring a Senior Platform Engineer to join the Platform engineering organization. The Platform team builds and operates the shared foundations that help Buildkite teams ship and scale products securely and reliably.
This is a hands-on role for a senior platform engineer with strong cloud and infrastructure foundations. You will design and operate reliable Kubernetes platforms across multiple clusters, solve distributed-systems challenges, automate infrastructure at scale, and take ownership of the systems you build.
The Platform team's mission is to build the foundations that let Buildkite teams ship, operate, and scale products securely and reliably. Buildkite's customers run some of the world's most demanding CI/CD and emerging AI workloads. As software development accelerates, CI/CD is becoming a major scaling bottleneck.
Current challenges include: self-service, repeatable foundations with secure defaults and reusable architecture; Kubernetes at production scale with improved networking, upgrades, patching, and capacity management; global and regional resilience with repeatable regional environments and disaster recovery; and safe, observable operations with improved observability, incident response, and runbooks.
Key responsibilities:
- Own substantial parts of Buildkite's AWS and multi-cluster Kubernetes platform from design through production operation
- Design reusable, self-service automation for provisioning and operating secure, consistent, and isolated platform environments
- Deliver regional architecture, request routing, disaster recovery, and tested failover capabilities
- Improve Kubernetes networking, upgrades, patching, scaling, observability, capacity management, and recovery
- Make infrastructure and deployments safer through secure defaults, progressive delivery, automated guardrails, and fast rollback
- Partner with product teams to diagnose distributed-system problems across service, datastore, network, and platform boundaries
- Turn incidents, capacity limits, and operational signals into durable engineering improvements
- Write well-tested, observable, documented, and operable systems; lead design discussions, review code, and mentor other engineers
A typical day might include designing reusable patterns for deploying or upgrading Kubernetes environments, working through how services operate safely across regional clusters, reviewing Terraform or architecture proposals, investigating capacity hotspots or network constraints with product teams, improving deployment guardrails and health signals, turning incident actions into automation, and pairing with or mentoring other engineers on reliability problems.
Buildkite is remote-first and async by default, with a focus on meaningful technical problems and visible impact. The role is 100% remote but must be located in either the ANZ or PST timezone.
REQUIREMENTS:
Core experience required:
- Clear communication and effective collaboration in a distributed team
- Technical influence: navigating ambiguity, explaining trade-offs, building consensus, and driving cross-team work
- Developer experience building CI/CD, internal-platform, or self-service capabilities that make the safe path easier for other engineers
- Experience owning production systems through growth, failures, migrations, and incidents
- Production AWS, Kubernetes, and infrastructure-as-code experience
- Infrastructure as Code: managing production infrastructure with Terraform or equivalent, including DNS, IAM, monitoring, and secrets
- Strong reliability practices, including observability, incident response, recovery testing, and operational improvement
- Experience building platforms or self-service capabilities that make secure, reliable delivery easier for engineers
Useful additional experience:
- Multi-region or multi-cluster systems, traffic routing, and disaster recovery
- Progressive delivery, deployment guardrails, and rollback
- Platform-level authentication, identity, or compliance controls implemented through engineering
- Datastores: operating relational databases, caches, or queues, including migrations, replication, failover, and workload isolation
Note: You do not need every item listed. The company is looking for demonstrated depth in several areas, sound judgement, and the ability to learn the rest.