SlipstreamJobsFresh Startup & VC-Backed Jobs

Data Engineer – Applied ML

SimilarWeb - Tel Aviv, Israel - Hybrid - posted 2026-09-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

SimilarWeb is a leading digital intelligence platform serving over 3,500 global customers including Google, eBay, and Adidas. The company went public on the NYSE in 2021 and continues to grow rapidly. You will join the R&D department as a Data Engineer with applied ML focus, working on the Retail Intelligence product line. Your mission is to transform raw product data collected from retailers and marketplaces worldwide into a single, trusted view by classifying products into unified taxonomies, normalizing brands and attributes, and matching entities across sources. This work operates at massive scale—over a billion records that grow and change daily. This is hands-on applied ML work, not pure research. You'll build production pipelines using LLMs, agentic frameworks (LangGraph), embeddings, and classical ML methods. Your responsibilities include: • Designing and building LLM-powered and ML-based pipelines for product, brand, and category data classification, normalization, structuring, and matching • Building agentic workflows (LangGraph or similar) to automate complex data tasks end-to-end • Selecting the right technical approach for each problem—LLMs, embeddings, fine-tuned models, classical classifiers, or rules—balancing accuracy, cost, and latency • Scaling solutions to run efficiently over billions of records using Spark, Databricks, and cloud infrastructure • Building evaluation frameworks: ground-truth datasets, labeling processes, quality metrics, and ongoing monitoring • Taking solutions from POC to production and owning them post-launch • Working closely with Product to define requirements and shape the roadmap • Collaborating with data engineers, data scientists, and other R&D teams on infrastructure and best practices The company emphasizes in-office collaboration with flexibility to work partially from home, fostering a connected, unified culture. **Requirements:** • B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics, or related field • 4+ years of hands-on experience as a data engineer, ML engineer, or data scientist with production solutions • Strong Python skills and production-quality code writing • Hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings, evaluation) • Experience with modern LLM stack: LLM provider APIs (OpenAI, Anthropic, etc.), LangGraph or LangChain, Hugging Face, and vector stores • Solid grasp of text classification and NLP methods (classical and modern) and judgment on when to use each • Experience processing large-scale data with Spark/PySpark, Databricks, or similar on AWS or other cloud platforms • Understanding of evaluation and data quality: precision/recall trade-offs, building ground truth, error analysis • Pragmatic, delivery-focused mindset; comfortable with ambiguity; clear communication with Product and business stakeholders • Advantage: experience with taxonomies, entity resolution, or product/e-commerce data • Advantage: experience with fine-tuning or deploying open-source models

Similar roles