SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 196,750 - 243,290 / annual
Roblox is a platform where tens of millions of people explore, create, play, and connect in 3D immersive digital experiences. The company operates thousands of microservices to support massive scale, and the Application Networking team is responsible for connecting and securing these services through ingress gateways, service mesh management, and seamless communication across hybrid on-prem and cloud infrastructure.
You will join the Service Mesh team to build the networking fabric that enables Cloud Bursting—a strategic initiative allowing Roblox's core services to transparently burst from on-prem data centers to the cloud, handling historical peak concurrent players and surviving regional failures.
Key responsibilities include:
- Design and build service mesh infrastructure enabling communication across hybrid Kubernetes and Nomad environments, supporting billions of daily requests
- Drive integration of service mesh with Kubernetes, ensuring reliable sidecar injection, mTLS, traffic policies, and observability for production workloads
- Build the networking foundation for Cloud Bursting with service discovery, traffic management, and locality-aware routing across environments
- Partner with internal application teams to understand connectivity pain points and provide a frictionless "paved road" for thousands of developers
- Collaborate with Gateway and CNI teams to deliver a unified, multi-cluster service fabric abstracting cluster and regional boundaries
- Serve as primary escalation point for complex service mesh issues, troubleshooting Envoy sidecars, Istio control plane, and underlying network stack
- Mentor junior engineers and promote best practices in testing, deployment, and reliability engineering
Requirements:
- 3+ years of experience in distributed systems with strong expertise in service mesh technologies (Istio, Envoy, Consul, or Linkerd); understanding of sidecar-based architecture tradeoffs at scale
- Deep knowledge of service mesh concepts: traffic management, service discovery, mTLS, observability, and routing policies
- Production Kubernetes experience and familiarity with K8s networking model
- Comfort with Envoy proxy internals, xDS APIs, and control plane architectures
- Ability to design large-scale distributed systems spanning multiple clusters, regions, and runtime environments
- Fluency in Go, C/C++, or Rust
- Experience with on-call rotations and reliability-first infrastructure mindset