SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 215,000 - 270,000 / annual
Starburst delivers enterprise intelligence at scale by giving organizations secure, governed access to all their data, wherever it lives. The company helps enterprises power AI and analytics without the cost and complexity of traditional data consolidation using open standards like Trino and Apache Iceberg.
You will join the AI layer team that builds AIDA, Starburst's AI agent product. This is the first dedicated research engineering hire on the team. You will own the intelligence layer that makes AIDA's agents correct, trustworthy, and measurably better over time. The work spans information retrieval, knowledge representation, and evaluation science.
Your core responsibilities include:
- Design and build grounding systems that connect agent reasoning to verified enterprise data sources
- Build and optimize retrieval pipelines (RAG, hybrid search, structured query generation) for accuracy and latency
- Define data representation strategies that preserve semantic fidelity across heterogeneous enterprise data (catalogs, schemas, lineage)
- Create evaluation frameworks: automated benchmarks, regression suites, human evaluation protocols
- Convert validated research findings into production systems that ship to users
- Establish quality metrics and dashboards that track agent correctness week over week
- Build feedback loops where user interaction data flows back into evaluation datasets and informs grounding improvements
You will operate at the research/systems boundary: running experiments with academic rigor and shipping results with production engineering discipline. Research and engineering are not separate tracks here. You will own experiments end to end, from hypothesis through production deployment. The team operates with startup speed inside an enterprise company, shipping weekly and measuring results.
REQUIREMENTS:
- 3+ years of experience in information retrieval, NLP, knowledge representation, or applied ML research
- Production experience building RAG, grounding, or retrieval systems (not prototypes or demos)
- Strong evaluation methodology: benchmark design, statistical analysis, reproducible experiments
- Comfort operating at the research/systems boundary: you read papers and you ship code
- Python fluency; experience with vector databases, embedding models, LLM APIs
- Track record of converting research insights into shipped production systems
PREFERRED QUALIFICATIONS:
- Experience with enterprise data systems (SQL engines, data catalogs, schema metadata)
- Familiarity with text-to-SQL or structured query generation
- Published research or open-source contributions in IR, NLP, or evaluation methodology
- Experience designing evaluation pipelines that run in CI/CD
- Familiarity with JVM-based systems
- Ability to travel 25% for onboarding, team offsites, customer engagements, and company events