SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Airbyte is the data and action layer for AI agents, providing fast, accurate, authenticated access to business data across hundreds of sources. The company has raised $181M from leading investors and operates at scale, processing over 3 million sync jobs per week across multiple regions and clouds.
You will be the infrastructure and reliability engineer on the Data Replication team, a full-stack product team responsible for the platform's infrastructure, reliability standards, incident reduction, and developer tooling. You'll work across Terraform files, Kubernetes clusters, and postmortem documentation.
Key responsibilities include:
- Own the infrastructure underpinning the Data Replication platform: Kubernetes clusters, CI/CD pipelines, secrets management, networking, and cloud resource configuration across AWS and GCP
- Partner with product engineers to reliably integrate product features with infrastructure
- Maintain and enhance observability, alerting, and anomaly detection with an eye towards LLM automation
- Maintain and enhance AI-augmented release and internal tooling: canary deployments, progressive rollouts, automated release qualification, and rollback automation
- Set the infrastructure bar for the team by building self-serve tooling, writing runbooks, and coaching engineers to own more of their stack
The role emphasizes active use of AI as a force multiplier—agentic tools to automate toil, augment incident response, and build smarter internal tooling. The company values trust, directness, and craftsmanship.
REQUIREMENTS:
- 7+ years in infrastructure, platform engineering, SRE, or DevOps
- Hands-on ownership of Kubernetes, Helm, and Terraform in production environments
- Deep experience with observability stacks (Prometheus, Grafana, Datadog) and on-call operations
- Experience with CI/CD pipeline ownership and developer tooling
- Ability and willingness to read backend code to understand how systems break and instrument them correctly
- Fluency with AI tools—LLMs and agentic frameworks to automate, debug faster, and reduce toil
- Startup-ready mindset: comfortable with ambiguity, moving fast, and owning problems end-to-end
NICE TO HAVE:
- Data pipelines, replication systems, or ETL/ELT platforms
- Control plane / data plane architectures or internal developer platforms
- Experience with Airbyte, CDKs, or connector-based architectures