SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 210,000 - 260,000 / annual
Blitzy is an AI software development platform that autonomously builds custom software for enterprises. The company routes enormous volumes of work across multiple LLM models, and the model landscape shifts constantly with new releases changing price-performance tradeoffs weekly.
You'll join as an early member of the applied AI research team, focused on model cost optimization. Your core mission: build the research, evaluation, and routing logic that lets Blitzy's platform dynamically select the most cost-effective model for every step of the software development lifecycle without sacrificing output quality.
Key responsibilities include:
- Design and maintain evaluation pipelines that benchmark models across cost, latency, and quality for Blitzy's core software development workflows
- Build and iterate on dynamic model-routing logic that selects the optimal model for a given task in real time
- Continuously track the model landscape and rapidly integrate and evaluate new model releases as they ship
- Partner with the platform engineering team to productionize routing and optimization logic at scale
- Define and track metrics that quantify the cost-to-quality tradeoff across the platform
Success means Blitzy's platform reliably selects the most cost-efficient model for each task without measurable drops in code quality, you've built a repeatable framework for benchmarking new models within days of release, model spend per unit of delivered work trends downward as usage scales, and engineering teams treat your evaluation infrastructure as the source of truth for model selection.
Required: strong applied machine learning or AI research experience with hands-on LLM evaluation/fine-tuning, solid software engineering skills to ship production code, familiarity with LLM evaluation methodologies and benchmarking frameworks, comfort with model inference economics (cost, latency, throughput), and a strong bias toward experimentation.
Standout qualifications: PhD in AI Systems or Applied AI Research, prior experience building model routing/selection/cascading systems in production, track record of publications or open-source contributions in LLM evaluation/optimization, experience with foundation model providers across multiple vendors, and founder's mentality.
This is a ground-floor opportunity to shape Blitzy's applied AI research strategy from day one. As an early team member, you'll have outsized influence over research direction, tooling, and infrastructure that the rest of the company will build on, with direct visibility into how your work affects both cost structure and platform capabilities.