SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 260,300 - 345,040 / annual
Roblox is seeking a Principal Platform Engineer to lead the ML Platform organization, which powers hundreds of use cases and billions of inferences daily across Discovery, Safety, Economy, and Creation. This role blends product thinking, developer experience, backend engineering, and infrastructure at scale.
You will own the ML Platform as a product end-to-end: define requirements, write RFDs, and ship APIs, SDKs, CLIs, and UIs that make machine learning adoption seamless across Roblox. You'll bootstrap and maintain core ML Platform components including the Serving Layer, Model Registry, Pipeline Orchestrator, and Training/Inference control planes. Your responsibilities include setting technical strategy and overseeing development of high-scale, reliable infrastructure systems with clear SLOs for latency, availability, and cost.
Key focus areas include designing exceptional developer experiences through paved-road templates, golden paths, opinionated defaults, and clear documentation to reduce time-to-first-production. You'll instrument the platform to measure adoption, friction, reliability, and cost, using data to prioritize roadmap and validate outcomes. Cross-organizational partnership is essential—you'll work with ML Engineering, Data Science, Infra/SRE, Security, and Finance to optimize performance, safety, and spend, particularly for GPU-intensive training and high-QPS inference.
Additionally, you'll propose and implement new platform tooling to improve time to production for ML engineers across the full ML lifecycle, stay abreast of industry trends in machine learning and infrastructure, mentor junior and senior engineers, lead design reviews, and drive cross-team architectural decisions.
Roblox operates on a hybrid schedule: Tuesday, Wednesday, and Thursday onsite at San Mateo headquarters, with optional presence on Monday and Friday.
REQUIREMENTS:
- 7+ years of professional experience with substantial system design expertise
- Proficiency in API design and developer experience (gRPC/REST APIs, SDKs, CLIs, UIs)
- Strong coding ability with demonstrated ability to ship product and obsess over user feedback
- Experience with end-to-end ML model lifecycle (model serving, training, model CI/CD, GPU resource management) preferred
- Passion for infrastructure-as-code and automating manual processes
- Ability to support ML engineers, understand their needs, and translate them into clean platform abstractions
- Excellent written and verbal communication across varying technical levels
- Bachelor's degree in Computer Science, Computer Engineering, Data Science, or similar technical field