SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 295,250 - 345,040 / annual
Roblox is building the future of human interaction through immersive 3D digital experiences. The Object Store team is developing an in-house, multi-tenant object storage platform that underpins all large-object workloads at Roblox—assets, media, analytics, backups, and machine learning datasets—operating at exabyte scale with trillions of objects and terabytes per second of throughput. The platform runs on bare metal in Roblox's own data centers, replacing cloud object storage with end-to-end control down to the hardware level.
As a Principal Engineer on the Object Store team, you will shape the architecture and set the technical direction for this critical infrastructure. Your responsibilities include:
• Design and build the object store components: S3-compatible gateway, metadata and indexing systems, data placement and tiering strategies, and multi-tenancy with isolation.
• Own the multi-year technical strategy, including build-vs-buy decisions and the migration off cloud object storage.
• Evolve the control plane to enable elastic scaling, autonomous healing, self-service provisioning, and zero-downtime tenant moves.
• Run the bare-metal fleet at scale, managing cluster lifecycle, upgrades, rebalancing, failure domains, and hardware/capacity planning.
• Profile and optimize critical data paths to drive down tail latency and cost per byte.
• Establish engineering best practices including design reviews, benchmarks, failure drills, and post-incident retrospectives targeting 4+ nines of reliability.
• Automate testing, CI/CD, rollout safety, observability, and autoscaling.
• Mentor engineers and partner cross-functionally with security, analytics, and product teams.
You will report to the Engineering Manager for Storage.
The role is based in San Mateo with hybrid flexibility: onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday.
**Requirements:**
• 8+ years of software engineering experience or equivalent.
• Deep experience building and operating large-scale distributed storage systems (object stores, blob stores, or distributed file systems) at petabyte scale or beyond.
• Hands-on production experience running an object storage platform: data placement and rebalancing, erasure coding, cluster tuning, and upgrades at scale. Operational depth is prioritized over internal contribution history.
• Strong proficiency in C++, Go, or Rust.
• Deep understanding of durability and replication: erasure coding, geo-replication, consistency trade-offs, and blast-radius containment.
• Proven track record shipping high-throughput, low-latency services on bare-metal, on-premises fleets with strong observability.
• Ability to translate ambiguous requirements into clear roadmaps and influence cross-functional stakeholders.
• Passion for automation, rigorous testing, and data-driven decision-making.
• Bonus: upstream contributions to open-source storage projects; S3 API compatibility at scale; NVMe/HDD tiering experience; or experience migrating large workloads off public-cloud object storage.