SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pulley is an AI-powered platform that streamlines commercial permitting across the US, helping architects, builders, and retailers accelerate project timelines. The company works with major brands like J.Crew, Solidcore, and Hibbett Sports, as well as data center and EV charging network projects. Founded in 2021 and backed by CRV, Susa Ventures, and Fifth Wall, Pulley combines deep permitting expertise with purpose-built AI.
In this Senior AI Engineer role, you will own AI-powered features end-to-end, from user research and problem definition through production deployment and iteration. The core technical challenge is transforming messy, unstructured inputs—scanned plan sets, jurisdiction codes, reviewer comments, and jurisdiction-specific application forms—into fast, structured, and trustworthy outputs.
Key responsibilities include:
- Owning AI features end-to-end: defining success criteria, designing prompts and pipelines, building evals, deploying to production, and iterating based on real-world performance
- Converting unstructured permitting documents and city regulations into reliable structured outputs using extraction, classification, retrieval, and agentic workflows
- Building evaluation and observability infrastructure for shipping LLM-powered features with confidence, including ground-truth definition, quality measurement, and regression detection
- Building with AI agents as a daily practice—directing, reviewing, and shipping agent-driven work while maintaining quality standards
- Making technical and product decisions with direct customer impact
- Raising the bar for engineers around you through design review, mentorship, and exemplary work
You thrive in ambiguity and are energized by undefined problems. You are product-minded, caring deeply whether your work solves real customer problems. You are rigorous about measurement—you trust evals over demos and build measurement before features. You hold strong opinions about balancing quality and velocity, and you default to ownership when something is broken or missing.
REQUIREMENTS:
- 4+ years of software engineering experience, with substantial time building production LLM or ML systems
- Track record of owning an LLM-powered product surface end-to-end, including data quality, eval design, cost, latency, and failure handling
- Deep hands-on experience with large language models in production: prompting, retrieval-augmented generation, structured extraction, tool use, agentic workflows, and knowing when each approach is inappropriate
- Experience designing evals and making LLM-powered features reliable in production
- Real experience building with AI coding agents—shipped work where agents performed substantial implementation under your direction
- Ability to architect durable systems while making pragmatic tradeoffs
- Based in San Francisco Bay Area and willing to work in-person 4 days per week
NICE TO HAVES:
- Document understanding at scale: OCR, layout-aware parsing, or vision-language models over scanned PDFs, drawings, or forms
- Fine-tuning models or building data pipelines for training and eval sets from real-world usage
- Experience in construction tech, govtech, proptech, or similar domains with messy real-world documents and processes
- Modern full-stack development with TypeScript, React, and Google Cloud
- Startup experience at the team-building stage
- Experience mentoring engineers or leading technical direction across teams