SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's ChatGPT Velocity team is hiring an experienced infrastructure engineer to dramatically improve developer experience and deployment reliability for ChatGPT-facing services. The team owns the end-to-end developer and deploy loop, from local environments through production, with a mission to help engineers ship confidently with fast, lightweight development cycles and trusted paths to production.
In this role, you will reduce friction across the entire development pipeline: local development startup and iteration latency, Bazel and build performance (cache effectiveness, reproducibility, disk usage, remote cache behavior), CI reliability and merge throughput (reducing unrelated failures, tail latency, flaky tests), and deploy pipeline reliability. You'll build agents to help PRs make forward progress through rebasing, conflict resolution, CI triage, and safe reruns. You'll also improve deploy observability so engineers can quickly understand where changes are, what's blocking them, and who owns the next action.
This is not a narrow tooling role. Success requires diagnosing messy distributed systems, reading developer pain signals from Slack and metrics, building reliable automation, negotiating ownership across teams, and making thoughtful tradeoffs between velocity and production safety. You should treat engineers as users and turn confusing workflow failures into clear, measured systems.
Key responsibilities include: reducing local development latency; improving Bazel and build performance across local machines, devboxes, and CI; improving CI reliability and merge throughput; building agents for PR progress (rebasing, conflict resolution, CI triage, safe reruns, owner routing); making deploy pipelines more self-healing; improving deploy observability; and establishing lightweight metrics for impact (commit-to-prod latency, deploy-window throughput, CI tail latency, local bootup time, failure rates, manual interventions, false-positive gates).
You should thrive in this role if you have strong experience with large-scale developer infrastructure, build systems, CI/CD, deployment systems, or production reliability. You can debug across local machines, devboxes, remote caches, monorepos, service dependencies, auth systems, and distributed deploy pipelines. You're comfortable with systems like Bazel, Buildkite, Kubernetes, Temporal, Python/TypeScript services, GitHub workflows, and large monorepo tooling, or can ramp quickly on equivalents. You think in product loops as much as systems loops, care whether developers understand and trust workflows, know how to replace manual intervention with policy-backed automation while preserving safety, and communicate well across teams. You prefer practical measurement over vibes while recognizing that developer frustration is often the first signal of a real systems problem. You're energized by high-leverage work where small improvements save hundreds of engineering hours.
Nice-to-have experience includes: improving developer experience in very large monorepos; remote execution, remote caching, hermetic builds, or cache reproducibility; designing CI quarantine, test ownership, merge queue, or auto-revert systems; progressive delivery, canary/preview/stable deploy pipelines, synthetics, alert quality, and deploy-manager workflows; building agent-assisted developer workflows (automated PR babysitting, CI triage, conflict resolution); and familiarity with security and access-control constraints around secrets, internal auth, VPN/exit-node networking, and local development.
Requirements:
- Strong experience with large-scale developer infrastructure, build systems, CI/CD, deployment systems, or production reliability
- Ability to debug across local machines, devboxes, remote caches, monorepos, service dependencies, auth systems, and distributed deploy pipelines
- Comfort with Bazel, Buildkite, Kubernetes, Temporal, Python/TypeScript services, GitHub workflows, and large monorepo tooling (or ability to ramp quickly on equivalents)
- Product-oriented thinking: care about developer understanding and trust, not just technical subsystem health
- Ability to replace repeated manual intervention with policy-backed automation while preserving production safety
- Strong cross-team communication skills; ability to turn ambiguous reports into crisp owners, hypotheses, metrics, and shipped fixes
- Preference for practical measurement over intuition, while recognizing developer frustration as a systems signal