SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, AI/ML Infrastructure

Thumbtack - ON, Canada - Remote - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 212,500 - 275,000 / annual

Thumbtack is seeking a Staff Software Engineer to lead the AI/ML Infrastructure team, which powers all AI-driven experiences across the platform—from search and recommendations to matchmaking, pricing, safety, content generation, and fraud detection. You will own technical direction across two critical surfaces: the AI/ML platform (LLM gateway, model training and serving, feature and data workflows, orchestration, observability) and the AI evaluation platform (trace ingestion, scorer infrastructure, LLM-as-judge calibration, pre-deploy regression gates, self-serve evaluation tooling). This is hands-on leadership: you design, build, and ship significant systems while elevating the technical bar of engineers around you. Key responsibilities include building scalable AI platform capabilities for GenAI applications, designing and deploying tools for applied scientists (model training/serving, feature workflows, CI/CD, orchestration, evaluation), working across the full stack from backend services to AI model integrations, evaluating next-generation AI infrastructure frameworks, and driving projects with measurable business impact. You will partner with applied scientists, product engineering teams, and leadership to define multi-quarter roadmaps and make consequential build/buy/adopt decisions in a fast-moving vendor and open-source landscape. Required: 10+ years software engineering with staff-level technical leadership scope; 3+ years building production AI/ML infrastructure (model serving, feature platforms, training, orchestration, observability); hands-on LLM production experience (gateways, prompt/model lifecycle, RAG, agentic architectures, cost optimization); experience building/operating AI evaluation systems (offline eval, LLM-as-judge calibration, regression testing, A/B measurement); deep expertise designing and operating reliable distributed systems with production ownership and on-call responsibility; demonstrated ability to use AI coding tools and validate/refine AI-generated output.

Similar roles