SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 88,000 - 176,000 / annual
Sourcegraph is building infrastructure for code understanding and agentic development. The company provides code search, deep search, and agentic batch changes that give engineering teams and AI agents cross-repository context to navigate massive codebases with confidence. Customers include Stripe, Reddit, and Leidos.
You will join the Code Understanding team as a staff-level ML and agentic systems engineer—a technical leader and production authority responsible for the hardest, most ambiguous problems in agent engineering. This is not a pure IC role; you will set technical direction, establish standards others follow, and be a force multiplier for a talented, product-minded team.
Key responsibilities:
**Agentic Systems**: Design and harden multi-step, tool-using agent loops that power current and new agentic experiences. Turn research and experiments into reliable, observable, and cost-bounded products at enterprise scale.
**Evaluations**: Bring judgment about where evaluations earn their keep versus where lightweight smoke tests suffice. Craft representative datasets, meaningful baselines, useful error taxonomies, and release criteria that connect offline measurements to production behavior. Avoid false rigor and move fast with confidence.
**Model Selection & Training**: Decide which models run where, drive upgrades, and fine-tune your own when warranted. Own the full lifecycle from dataset construction through evaluation, production rollout, and monitoring.
**Retrieval & Context Engineering**: Push on how the system grounds models in customer code—retrieval, ranking, context windows, citations—to make answers more accurate and verifiable.
**Cost & Latency**: Treat cost and latency as product features. Profile, distill, cache, and right-size models so ambitious features ship sustainably. Make measured quality-latency-cost tradeoffs using the right combination of techniques rather than reflexively reaching for more complex models.
**Mentorship & Technical Leadership**: Up-level teammates through pairing on hard problems, substantive design and code reviews, and spreading agent engineering literacy. Operate autonomously on ambiguous problems, own high-technical-risk projects end-to-end, and contribute beyond your immediate domain.
Within one month, you will get the Code Understanding products running locally and land your first improvements to a model, prompt, retrieval path, or eval. Within three months, you will own a meaningful agentic slice of the product end-to-end and establish how the team ships model and prompt changes responsibly. Within six months, you will be the recognized technical authority for agent engineering on Code Understanding and measurably move the products in quality, cost, latency, or new capabilities.
The team is small, senior-leaning, and ships quickly. Engineers talk to customers, frame problems, and own them end-to-end. You will have real agency over technical direction and a direct line to the impact of your work.
**Requirements:**
- You have personally owned a production model lifecycle: trained or fine-tuned at least one model and taken it from dataset construction through evaluation, production rollout, and monitoring. You can explain how you chose between training/fine-tuning and prompting/retrieval, how you chose baselines and metrics, what error analysis revealed, and how production evidence affected the next version.
- You build agents fluently and opinionatedly: designed multi-step agentic systems and made them reliable, observable, and cost-bounded. You have a point of view on where agents shine and where deterministic code or human judgment is required.
- You have strong evaluation judgment: build representative datasets, meaningful baselines, useful error taxonomies, and release criteria that connect offline measurements to production behavior. You know where rigorous evaluations earn their keep, where lightweight smoke tests or qualitative review are enough, and where a precise-looking metric is misleading.
- You treat cost and latency as product constraints: make measured quality, latency, and cost tradeoffs and use the appropriate combination of model selection, prompting, retrieval, caching, distillation, and fine-tuning.
- You operate autonomously on ambiguous problems: given a rough product idea, customer quotes, and a Slack thread, you come back with a plan, prototype, milestones, and a point of view on tradeoffs without pre-scoping. You own high-technical-risk projects end-to-end.
- You contribute beyond your domain: go into whatever part of the codebase a problem requires, recognize issues beyond your immediate area, and translate between engineering goals and business objectives.
- You up-level people around you: mentor by pairing on hard problems, providing substantive design and code reviews, and spreading agent engineering literacy. You see investing in teammates' growth as part of the job.
- Working hours must overlap with EST for at least 20 hours per week. Preferred locations are Europe or North America, though applications from anywhere are welcome.