SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 184,000 - 314,000 / annual
Braze is seeking a Staff Machine Learning Engineer to join the Predictive and Generative AI (PGAI) team. The team's mission is to deliver engaging and personalized customer experiences through ML and AI-enhanced marketing solutions, operating those systems at global scale with distributed training pipelines for customer-specific models and high-throughput APIs serving predictions across multiple regions.
In this hands-on Staff role, you will own the ML platform infrastructure and drive transformative initiatives that change how the team runs ML in production. Key responsibilities include:
• Identify and drive major platform initiatives such as replatforming queueing and orchestration systems, overhauling deployment and cloud identity, or retiring infrastructure generations
• Build and ship at high velocity, carrying the most complex infrastructure initiatives from design through production, including multi-region model serving fleets, customer-specific model health pipelines, and CI/deployment tooling
• Own the platform's technical vision and production quality bar, setting direction for model training, deployment, serving, and observability; lead incident response for ML systems; drive reliability and cost optimization at scale
• Drive cross-team initiatives leveraging shared infrastructure, deployment tooling, and data systems owned with partner teams, maintaining technical relationships across boundaries
• Raise engineering quality through design review, code review, and production readiness assessment for ML systems; mentor other senior engineers and data scientists
• Connect technical decisions to customer and business outcomes, representing the team's technical perspective to product and engineering leadership
This is a hands-on delivery role where you will personally own and ship complex infrastructure work while providing technical leadership and mentorship.
REQUIREMENTS:
• 8+ years building and operating distributed systems in production, with depth in deployment and operations; experience designing services for scale and reliability, owning CI/CD and infrastructure as code, and running systems under production load
• Hands-on experience with ML workloads in production (training pipelines, model serving, feature systems, or ML platform tooling); deep modeling experience is a plus rather than a requirement
• Technical leader who has owned direction for a team, led multi-quarter initiatives across team boundaries, and grown senior engineers while maintaining high personal output
• Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management, networking, and cost profiling
• Effective communicator (verbal and written) whose designs and recommendations build consensus and drive decision-making
BONUS QUALIFICATIONS:
• Experience with queueing and orchestration systems (Celery, RabbitMQ, Kafka, Ray)
• ML platform tooling experience (MLflow, model registries, feature stores, ML observability)
• Familiarity with Braze's stack (Python, Ruby on Rails, MongoDB, Redis, Kubernetes)
• Operating under compliance regimes (SOX, HIPAA)
• Customer engagement, personalization, or marketing technology domain experience