SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pulley is an AI-powered platform that streamlines commercial permitting across 19,000+ jurisdictions. The company helps architects, builders, and retailers accelerate project timelines by eliminating costly delays in the permitting process—historically the slowest and most uncertain phase of building. Pulley serves major brands like J.Crew, Solidcore, and Hibbett Sports, as well as data center buildouts and EV charging networks.
As a Staff AI Engineer, you will own the AI problem space across the product, not just individual features. Your core responsibilities include:
- Define technical direction for how Pulley applies LLMs across multiple product surfaces, from ambiguous problem definition through architecture to shipped, iterated-on product.
- Transform unstructured permitting documents, city regulations, and jurisdiction workflows into structured, reliable outputs using extraction, classification, retrieval, and agentic workflows.
- Establish evaluation and observability standards company-wide: define ground truth, measure quality and regressions, and build systems that make rigorous measurement the default for every team shipping LLM features.
- Build and ship agent-driven work at high velocity while owning the quality bar; direct and review AI agents as a daily practice.
- Make technical bets that determine Pulley's capabilities next year—which models, architectures, and build-versus-buy decisions—and own the production consequences.
- Multiply the engineers around you by setting patterns for LLM feature development, mentoring senior engineers toward larger scope, and accelerating the team through systems, standards, and abstractions.
You thrive in ambiguity and are energized by undefined problems. You are product-minded and care deeply whether your work actually solves customer problems. You are rigorous about measurement and don't trust demos—you trust evals. You hold strong opinions about quality and velocity as complementary, not competing, goals. You default to ownership at organizational scale and take initiative to fix broken systems or processes without waiting for permission.
REQUIREMENTS:
- 8+ years of software engineering experience, with substantial time building production LLM or ML systems.
- Track record of owning a significant AI domain end-to-end: from problem identification through architecture, delivery, and production ownership, including data quality, eval design, cost, latency, and failure handling.
- Deep hands-on production experience with large language models: prompting, retrieval-augmented generation, structured extraction, tool use, agentic workflows, and knowing when each approach is wrong.
- Experience designing evals and making LLM-powered features reliable in production.
- Real experience building with AI coding agents—not just autocomplete; you have shipped work where agents performed substantial implementation under your direction.
- Ability to architect durable systems while making pragmatic tradeoffs.
- Experience mentoring engineers or setting technical direction that others built within.
- Based in the San Francisco Bay Area and willing to work in person 4 days per week.
NICE TO HAVES:
- Document understanding at scale: OCR, layout-aware parsing, or vision-language models over scanned PDFs, drawings, or forms.
- Model fine-tuning or building data pipelines to produce training and eval sets from real-world usage.
- Domain experience in construction tech, govtech, proptech, or similar fields with messy real-world documents and processes.
- Modern full-stack development: TypeScript, React, Google Cloud, and comfort working in application code that surfaces AI features to users.
- Experience as the most senior AI engineer in a domain—the person others escalated to when nobody knew the answer.
- Startup experience at the stage where you helped build the team, not just the product.