SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Machine Learning Engineer, ML Platform

Braze - Austin, TX, United States - Hybrid - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 184,000 - 314,000 / annual

Braze is seeking a Staff Machine Learning Engineer to join the Predictive and Generative AI (PGAI) team. The team's mission is to deliver engaging and personalized customer experiences through ML and AI-enhanced marketing solutions, operating those systems at global scale with distributed training pipelines and high-throughput prediction APIs across multiple regions. In this role, you will own the ML platform infrastructure and make deploying, operating, and scaling ML at Braze fast, safe, and efficient. Key responsibilities include: - Identify and drive transformative initiatives that change how the team runs ML in production, such as replatforming queueing and orchestration, overhauling deployment and cloud identity, or retiring infrastructure generations. - Build and ship at high velocity as a hands-on delivery role, carrying the most complex infrastructure initiatives from design through production. Current examples include multi-region model serving fleets, pipelines maintaining hundreds of customer-specific models, and CI/deployment tooling. - Own the platform's technical vision and production quality bar. Set direction for model training, deployment, serving, and observability; lead incident response for ML systems; and drive reliability and cost optimization at scale. - Drive cross-team initiatives. The platform builds on shared infrastructure and data systems owned with partner teams; you will carry technical relationships with those teams. - Raise engineering quality through design review, code review, and production readiness for ML systems. Mentor other senior engineers and data scientists. - Connect technical decisions to customer and business outcomes, representing the team's technical perspective to product and engineering leadership. Braze is a leading customer engagement platform that empowers brands to deliver personalized experiences. The company is headquartered in New York with offices globally including Austin, Berlin, Chicago, London, San Francisco, São Paulo, Singapore, Sydney, and Tokyo. REQUIREMENTS: - 8+ years building and operating distributed systems in production, with depth in deployment and operations. Must have designed services for scale and reliability, owned CI/CD and infrastructure as code, and run systems under production load. - Hands-on experience with ML workloads in production: training pipelines, model serving, feature systems, or ML platform tooling. Deep modeling experience is a plus but not required. - Technical leader who has owned direction for a team, led multi-quarter initiatives across team boundaries, and grown senior engineers while maintaining high personal output. - Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management, networking, and cost profiles. - Effective communicator (verbal and written) whose designs and recommendations build consensus and drive decision-making. BONUS QUALIFICATIONS: - Experience with queueing and orchestration systems (Celery, RabbitMQ, Kafka, Ray). - ML platform tooling experience (MLflow, model registries, feature stores, ML observability). - Familiarity with Braze's stack (Python, Ruby on Rails, MongoDB, Redis, Kubernetes). - Experience operating under compliance regimes (SOX, HIPAA). - Customer engagement, personalization, or marketing technology domain experience.

Similar roles