SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 212,500 - 323,400 / annual
Thumbtack is seeking a Staff Software Engineer to lead the AI/ML Infrastructure team, which powers all AI-driven experiences across the platform—from search and recommendations to matchmaking, pricing, safety, content generation, and fraud detection.
You will own technical direction across two critical surfaces: the AI/ML platform (LLM gateway, model training and serving, feature and data workflows, orchestration, observability) and the AI evaluation platform (trace ingestion, scorer infrastructure, LLM-as-judge calibration, pre-deploy regression gates, self-serve evaluation tooling).
This is a hands-on role where you design, build, and ship significant systems yourself while elevating the technical bar of engineers around you. You'll partner with applied scientists, product engineering teams, and leadership to define multi-quarter roadmaps, make build/buy/adopt decisions across a fast-moving vendor and open-source landscape, and ensure the platform scales with Thumbtack's rapidly growing AI ambitions.
Key responsibilities include building core AI platform capabilities for GenAI-powered applications, designing scalable tools for model training, serving, feature workflows, CI/CD, orchestration, and evaluation. You'll work hands-on across the stack—backend services, execution infrastructure, and AI model integrations—while evaluating next-generation frameworks and driving projects with measurable business impact.
Required: 10+ years professional software engineering with staff-level technical leadership scope; 3+ years building production AI/ML infrastructure (model serving, feature platforms, training, orchestration, observability); hands-on LLM production experience (gateways, prompt/model lifecycle, RAG, agentic architectures, cost optimization); experience building AI evaluation systems (offline eval, LLM-as-judge calibration, regression testing, A/B measurement); deep expertise designing and operating reliable distributed systems at scale with on-call and incident leadership; demonstrated proficiency using AI coding tools and validating AI-generated output.