SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Similarweb is a leading digital intelligence platform serving over 3,500 global customers including Google, eBay, and Adidas. The company went public on the NYSE in 2021 and continues to grow rapidly.
You will join the R&D department as a Data Engineer with applied ML focus, working on Retail Intelligence products. Your mission is to transform raw data collected from retailers and marketplaces worldwide into a single, trusted view by classifying products into unified taxonomies, normalizing brands and attributes, and matching entities across sources. This work operates at scale across a catalog of over one billion records that continuously grows and changes.
This is a hands-on applied ML role within data engineering—not pure research. You will build production pipelines using LLMs, agentic frameworks (LangGraph), embeddings, and classical ML methods. Success requires deep understanding of classification and NLP methods to select the right tool for each problem and validate results.
Key responsibilities:
- Design and build LLM-powered and ML-based pipelines for classifying, normalizing, structuring, and matching product, brand, and category data
- Build agentic workflows (LangGraph or similar) that automate complex data tasks end-to-end
- Select the optimal approach for each problem (LLMs, embeddings, fine-tuned models, classical classifiers, or rules), balancing accuracy, cost, and latency
- Scale solutions to run efficiently over billions of records using Spark, Databricks, and cloud infrastructure
- Build evaluation frameworks including ground-truth datasets, labeling processes, quality metrics, and ongoing monitoring
- Take solutions from POC to production and own them post-launch
- Work closely with Product to define requirements and shape the roadmap
- Collaborate with data engineers, data scientists, and other R&D teams on infrastructure and best practices
Requirements:
- B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics, or related field
- 4+ years of hands-on experience as a data engineer, ML engineer, or data scientist with production solutions
- Strong Python skills and production-quality code writing
- Hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings, evaluation)
- Experience with modern LLM stack: LLM provider APIs (OpenAI, Anthropic, etc.), LangGraph or LangChain, Hugging Face, vector stores
- Solid grasp of text classification and NLP methods (classical and modern) with knowledge of when to apply each
- Experience processing large-scale data with Spark/PySpark, Databricks, or similar on AWS or other cloud platforms
- Understanding of evaluation and data quality: precision/recall trade-offs, building ground truth, error analysis
- Pragmatic, delivery-focused mindset; comfortable with ambiguity; clear communication with Product and business stakeholders
- Advantageous: experience with taxonomies, entity resolution, or product/e-commerce data; experience fine-tuning or deploying open-source models