SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 388,000 - 619,000 / annual
Netflix Cloud is the infrastructure backbone powering every Netflix service. The Infrastructure Management organization builds the platforms that enable engineers across the company to provision, configure, and operate infrastructure safely and at scale. This role owns critical components: Netflix's infrastructure-as-code product (the primary interface to Netflix Cloud, integrated with Spinnaker and the SDLC), a durable execution platform powering Tier-0 control planes, and emerging agent sandbox primitives for safe AI/agentic workload execution across compute environments.
As a Staff Engineer, you will drive high-leverage product decisions that scale developer platforms from early adoption to exponential usage across the entire engineering organization. You own the end-to-end Netflix Cloud developer experience, spanning discovery, configuration, deployment, operations, and debugging. You build and operate Tier-0 control planes with direct accountability for reliability, scalability, and cost. You define how AI agents safely interact with Netflix infrastructure by designing sandboxing, identity, and governance primitives, plus shared human/agent interfaces. You set critical boundaries between self-serve capabilities and guardrails by choosing platform defaults, provider guarantees, and areas where the platform intentionally steps back.
You think like a startup founder, partner deeply across engineering organizations, and treat the platform as a product with real users and real stakes. You have built or meaningfully scaled developer platforms in production (not just contributed), consistently make strong product decisions by building backward from customer goals and adoption constraints, and are comfortable operating in ambiguity—setting direction with minimal structure, iterating via PoCs, and converging quickly on simpler solutions. You are fluent in multi-language, multi-service codebases (Java/Kotlin/Go or equivalent), ramp quickly on unfamiliar systems, and have strong distributed systems fundamentals including race conditions, signal synchronization, and control plane design—with production failures to prove your depth.