SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 145,000 - 185,000 / annual
Traba is an AI operating layer for the industrial supply chain, starting with workforce management (temp staffing) and expanding into broader operational workflows. The company has proprietary data from millions of shifts, deep enterprise relationships, and backing from Founders Fund, Khosla Ventures, and General Catalyst.
As Staff Data Scientist, you'll lead measurement and modeling for Traba's agentic platform from inception. You'll own how the company models and understands agent performance in real customer workflows—measuring capability, reliability, and unit economics—and partner with engineering, product, and operations leadership on platform-shaping decisions.
Key responsibilities include: providing strategic insights to senior leadership through statistical analysis and modeling; designing and maintaining metrics, models, and reporting for agent quality, reliability, adoption, and unit economics; building evaluation and experimentation as a first-class discipline with production datasets, rubrics, automated graders, and regression suites; identifying business challenges like agent failure modes and cost/latency patterns, then building statistical and ML models to drive improvements; architecting scalable analytics and modeling infrastructure with proper data governance; overseeing the data warehouse; working closely with the Agents team and Operations leadership to understand data needs and provide actionable insights; enabling Operations teams with self-service analytics and AI-assisted tools; and mentoring scientists and analysts while setting standards for measurement, modeling, and experimentation.
Required: 7+ years in data science, ML, applied statistics, or quantitative research, with 2+ years hands-on modeling or measuring LLM/agent-based systems in production. BS/MS/PhD in data science, statistics, ML, computer science, mathematics, economics, or related quantitative field (or equivalent). Strong Python proficiency (scikit-learn, PyTorch, pandas, statsmodels), SQL expertise, A/B testing and statistical inference experience, and familiarity with LLM evaluation tools (Langfuse, Braintrust). Excellent data storytelling, cross-departmental collaboration, curiosity, and ability to work independently in a fast-paced startup.
Bonus: Jupyter/Hex/Hyperquery, dbt, internal agents/MCP servers, vertical AI experience, LLM fine-tuning/distillation, or causal inference at scale.