SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mapbox is the leading real-time location platform serving over 4 million registered developers. This Senior AI Engineer role is scoped by technical expertise rather than a single product, spanning Search and Places data, Location and Navigation Intelligence, and Platform work that enables AI agents to use Mapbox.
You will own the technical design and delivery of multi-component AI systems, accountable for quality and production performance. Key responsibilities include:
- Define product behavior, build measurement frameworks, and create evaluation systems for non-deterministic AI outputs. Formulate and validate hypotheses through appropriate datasets and regression testing.
- Run continuous evaluation of APIs, SDKs, data representations, and reference applications from end-user perspectives (developers, agents, consumers). Identify integration gaps and recommend fixes.
- Own MVP delivery against agreed technical designs, balancing perfection with shipping useful increments.
- Build and maintain data pipelines: ingestion, conflation, entity resolution, quality checks, batch and streaming jobs for large datasets.
- Track external datasets, models, and benchmarks from research and open-source. Decide when to adopt versus build.
- Design feedback loops where product usage generates improvement data. Instrument systems for reproducible failures.
- Design boundaries between models and their tool calls. Build model harnesses and manage delegation and state management.
- Optimize for latency and cost per request through streaming, caching, model routing, and prompt structure.
- Build internal tools (CLI, MCP, etc.) for team iteration and share generalizable components.
- Raise team standards through code review and evaluation practice mentorship.
- Participate in on-call rotation for 24/7 system availability, including potential off-hours response.
This is a distributed, asynchronous-first organization where most decisions happen in documents and Slack.
REQUIREMENTS:
- Bachelor's degree in STEM discipline
- 5+ years software engineering experience with production ownership of services, pipelines, or SDKs
- 2+ years shipping LLM-backed features to real users in systems with error budgets, on-call rotations, and customer-facing regressions
- Direct experience or deep understanding of evaluation design for non-deterministic systems; ability to describe a dataset built and failure caught
- Data engineering depth: SQL, at least one distributed processing framework, experience with pipelines where data quality matters more than speed
- Fluency with tool calling and agent orchestration, including failure modes (stale context, hallucinated arguments, silent partial success, unbounded loops)
- Working knowledge of multiple agent harnesses with informed opinions on strengths/weaknesses
- Strong Python or TypeScript; comfort reading code in any language
- Experience diagnosing latency in distributed request paths
- Comfort with ambiguity and judgment to ship narrow working solutions while general solutions remain unclear
- Clear written communication for distributed, async-first environment
NICE TO HAVE:
- Geospatial data experience (routing, geocoding, POI/address data, OpenStreetMap, conflation)
- Public API or SDK design for external developers
- MCP or similar tool transport experience
- CI-based eval running with commercial or custom harnesses
- Automotive/in-vehicle infotainment/CarPlay/Android Auto experience
- Voice pipeline experience (streaming ASR, TTS, barge-in, endpointing, wake word)
- Constrained compute, offline, or intermittent connectivity experience
- Product launch experience through first external integrations