SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
GoFundMe is seeking a Manager of Machine Learning Engineering to lead the team responsible for ML/AI infrastructure, pipelines, and operational excellence across the platform. In this role, you will own the reliability, scalability, and health of production ML/AI systems—including training pipelines, feature stores, model serving, and monitoring infrastructure.
You will lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and best practices for production systems. Key responsibilities include partnering with data science and ML engineering teams to streamline the path from model development to production deployment; establishing org-wide ML operational excellence through standards for model observability (latency, errors, drift, calibration, business KPI deltas); and building mature on-call processes, SLOs/SLAs, and postmortem practices that treat model incidents with the same discipline as infrastructure incidents.
You will drive operational strategy for both traditional ML and generative AI systems, balancing innovation velocity with safety, compliance, cost, and reliability. You'll collaborate cross-functionally with Product, Engineering, Design, and Legal/Privacy to translate business goals into team priorities and measurable outcomes. You'll also manage vendor and platform relationships (cloud ML platforms, LLM providers) and make build-vs-buy decisions.
Required: 7+ years hands-on experience building and shipping production ML systems in high-availability environments; 1-3+ years directly managing engineers in MLOps, ML platform, or infrastructure contexts; strong Python proficiency and ML frameworks (PyTorch, TensorFlow, Scikit-learn); experience designing and operating real-time model serving at scale; strong data engineering fluency with SQL, Spark/Databricks, and warehouse technologies; proven ML monitoring implementation for technical and business metrics; and strong leadership and mentoring skills.
Preferred: Familiarity with generative AI/LLM infrastructure and operational considerations; advanced degree in Computer Science, Statistics, Data Science, or related field. In-office requirement: 3 days per week in San Francisco Bay Area.