SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Monday.com is an AI work platform serving 250,000+ customers. The AI Infra group builds foundational infrastructure, tools, and platforms that enable intelligent, agentic features across the organization. This includes the AI Gateway and a centralized evaluation framework ensuring all AI features deployed to production are secure, resilient, cost-effective, and trustworthy.
As a Data Scientist on the AI Infra team, you will bridge AI research and production infrastructure, partnering with engineering and product teams to design evaluation judges, metrics, and error analysis workflows. Your work will enable monday to ship cutting-edge AI agents with speed and confidence.
Key responsibilities:
- Own evaluation methodology: Design metrics, pipelines, and evaluation approaches that teams across monday adopt as their source of truth for measuring AI quality.
- Transform the AI agent lifecycle: Standardize how AI agents are built, regression-tested, and maintained across the organization, embedding continuous evaluation into engineering workflows and post-deployment monitoring.
- Drive organizational impact: Partner with AI feature teams to translate domain expectations into meaningful datasets, test suites, and continuous evaluation pipelines, while leveling up engineers and product managers.
- Build hands-on tools: Prototype and deploy evaluation pipelines end-to-end, converting high-level product requirements into clear, quantifiable evaluation standards that gate production releases.
- Anticipate failure modes: Proactively design next-generation evaluation strategies for evolving agent architectures.
You will work at monday.com's Tel Aviv headquarters in an ownership-driven culture where you shape how organizations work.
Requirements:
- 3+ years of experience as a Data Scientist in non-academic, production settings working with complex AI systems. Familiarity with agentic frameworks like LangGraph, LangChain, and state-of-the-art SDKs.
- Deep, practical understanding of how agents operate—models, context, capabilities, and harnesses. Deep experience with agentic evaluation methodologies.
- Production-grade coding skills with a track record of building, prototyping, and shipping end-to-end data or evaluation pipelines.
- "Trace-first" diagnostic mindset: comfortable diving into raw agent execution logs, inspecting failure modes, and constructing qualitative error taxonomies.
- Strong product intuition and exceptional communication skills to translate complex evaluation data into clear, actionable guidelines.
- Proven ability to partner closely with software engineers and product teams, driving adoption and raising quality standards across the organization.
Preferred qualifications:
- Direct experience designing evaluation strategies for complex agentic systems in production.
- Prior experience within centralized platform/infra teams supporting multiple product verticals.
- Experience with modern microservice architectures, GitHub workflows, and automated production CI/CD pipelines.
- Familiarity with modern agent and evaluation tooling and observability stacks (e.g., LangSmith, Langfuse, or custom internal platforms).
- Knowledge of TypeScript or experience with modern platform architectures.