SlipstreamJobsFresh Startup & VC-Backed Jobs

Product Analyst

Bolna AI - Bengaluru, India - In-office - posted 2026-09-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Bolna is a YC-backed voice AI orchestration platform built for the Indian market, powering multilingual, vernacular voice agents across Hindi, Hinglish, Tamil, and 10+ languages at sub-500ms latency. The company serves collections, recruitment, sales, and e-commerce use cases. This Product Analyst role bridges model/evaluation rigor and product/growth insight—two areas currently owned by the Head of Product. You will own the execution and recurring cadence of data analysis across both domains, freeing product leadership to act on findings rather than produce them. KEY RESPONSIBILITIES: Model and Evaluation Analysis: - Run structured LLM and model benchmarking across providers (Sarvam, DeepSeek, Gemini, Claude) for post-call extraction and LLM-as-judge scoring, evaluating cost, accuracy, fill rate, and time-to-response with attention to Hinglish and code-mixed content. - Build and maintain LLM-as-judge pipelines using tools like DeepEval; design and track evaluation metrics; run inter-rater reliability analysis (e.g., Krippendorff's alpha) across human call reviewers. - Support golden dataset construction for ASR and transcript labelling, flagging conventions around code-switching in Devanagari vs. Roman script and transliteration normalization. - Evaluate WER and voice quality metrics across ASR providers for Indic languages using public benchmarks and academic references. - Analyse routing, latency, and cost data (including Azure PTU utilization and percentile latency distributions) to inform infrastructure decisions. - Support population-level analysis of graph-agent behaviour, including node-level aggregates, designed-vs.-observed differences, stuck-in-loop detection, and failure-pattern metrics. Product and Growth Insight: - Instrument and analyse the self-serve funnel (signup → activation → habit → expansion); identify drop-off points and friction. - Mine aggregate call data across customers and use cases for patterns: completion rates by use case, best-performing model/configuration combinations, emerging failure patterns in prompt templates. - Serve as a fast, reliable "pull me the data on X" resource for pod PMs, covering usage patterns, cohort behaviour, and feature adoption. Reporting and Tooling: - Build repeatable dashboards and scripts (not one-off notebooks) so analyses run on an ongoing cadence. - Present findings to product, ML, and infrastructure stakeholders in actionable form. FIRST 90 DAYS SUCCESS METRICS: - Execute end-to-end LLM benchmarking for post-call extraction, producing clean cost, accuracy, and fill-rate tables. - Contribute meaningfully to golden dataset build by flagging transliteration and script inconsistencies. - Run inter-rater reliability pipeline on a recurring basis. - Produce a clear first-pass self-serve funnel view with at least one concrete drop-off point identified. - Become the reliable first pass for "pull me the data on X" across at least two analytical areas. REQUIREMENTS: Must-Have: - 0–2 years of experience as a new graduate or early-career professional in data analysis, product analytics, applied ML, or research-adjacent roles. Hiring for raw analytical strength and trainability. - Strong SQL and Python skills (pandas); comfortable writing and debugging queries and scripts independently. - Solid statistical fundamentals: distributions, basic hypothesis testing, agreement/reliability concepts. Ability to learn new statistical concepts quickly. - Genuine comfort with ambiguous, messy real-world data; ability to notice when something looks wrong and flag it clearly. - Basic funnel and cohort analysis instincts; comfort thinking in terms of drop-off stages and segments. - Native or near-native fluency in Hindi or another Indian language, plus English; comfort reading and labelling code-mixed or Hinglish text. Strong Plus: - Exposure to LLM evaluation concepts (prompt-based scoring, LLM-as-judge), speech/ASR evaluation (WER, transcript QA), or evaluation frameworks (DeepEval) through coursework or personal projects. - Exposure to product or growth analytics (funnel analysis, retention curves, cohort behaviour) via prior role, internship, or personal project. - Familiarity with cloud-inference economics (token-based billing, provisioned throughput models). - Portfolio of self-directed analysis (project, competition, write-up) demonstrating ability to identify the "so what," not just the number.

Similar roles