Our mission is to bring clarity and control to the world's most complex codebases. AI is accelerating code creation, but the infrastructure to understand, oversee, and evolve that code hasn't kept pace. Sourcegraph gives engineering organizations full visibility across their systems, precise context for their agents, and the ability to execute coordinated code changes at scale.With Code Search, Deep Search, MCP, and Agentic Batch Changes, we deliver on that mission today - giving engineering teams and their AI tools the cross-repo context to navigate massive codebases with confidence, and the ability to make changes across hundreds of repositories at once.As the staff ML and agent engineer on Code Understanding, you'll be the technical owner of it: setting the tam’s direction for models, evaluations, and agentic systems, making our products measurably better, faster, and cheaper, and raising the team's fluency in building with models.This is a staff-level role: we're hiring a technical leader, not just a strong individual contributor. The primary need is a production ML and evaluation authority who also builds production agent systems. You'll own the hardest, most ambiguous problems in this space, set standards others follow, and influence direction beyond your immediate team.Concretely, you'll own work like:Agentic systems. You'll design and harden the multi-step, tool-using agent loops behind current and new agentic experiences, turning research and experiments into reliable, observable, and affordable products at enterprise scale.Pragmatic use of evaluations. Crafting agentic products means needing to tell when a change actually helped, which is hard when agents keep changing and the product keeps shifting. You'll bring judgment about where evaluations earn their keep, when to use targeted smoke tests and metrics, and how to avoid noise dressed up as rigor, so we can move fast with confidence.Models: selection, upgrading, and training. You'll decide which models we run where, drive upgrades, and fine-tune our own when that's the right call.Retrieval and context engineering. You'll push on how we ground models in a customer's code - retrieval, ranking, context windows, citations - to make answers more accurate and verifiable.Cost and latency. Every surface has a per-user economic budget. You'll treat cost and latency as product features, and profile, distill, cache, and right-size models so we can ship ambitious features sustainably.You'll do this on a small, senior-leaning team that ships quickly, owns a lot of product surface, and has streamlined product management: engineers here talk to customers, frame the problem, and own it end-to-end. You'll have real agency over technical direction and a direct line to the impact of your work.📅 Within one month, you will…Get the Code Understanding products and their model/agent pipelines running end-to-end locally, and land your first improvements to a model, prompt, retrieval path, or eval.Build a clear picture of where the AI engineering pain is, and the product surfaces most constrained by them.Get to know the team and our customers, and start forming your own opinions about where our agentic products should go next.Join the team's on-call support rotation.📅 Within three months, you will…Own a meaningful agentic slice of the product end-to-end, driving it from problem-framing through rollout and measurement.Establish how the team ships model and prompt changes responsibly: the evals, dashboards, and guardrails that make quality and cost regressions visible before customers feel them.Begin up-leveling teammates in building with models by pairing with them, reviewing their code, and modeling good agent engineering instincts.📅 Within six months, you will…Be the recognized technical authority for agent engineering and agentic systems on Code Understanding. You will be the person teammates, and increasingly the wider department, defer to on model, eval, and agent-design decisions.Have measurably moved the products: better answer quality, lower cost/latency, or new agentic capabilities that weren't feasible before.Be setting the direction of the team's roadmap where it intersects agents, bringing conviction, backed by evidence, about which bets are worth making, and pulling other engineers up to execute on them.About you You are a staff engineer and technical leader with hard-won skills across production machine learning, evaluation, and agent systems. This high-leverage role relies on your ability to make sound model and evaluation decisions for a fast-moving product, build the production systems around them, steer technical direction, and be a force multiplier for a talented, product-minded team.You have personally owned a production model lifecycle. You have trained or fine-tuned at least one model and taken it from dataset construction through evaluation, production rollout, and monitoring.You build agents, fluently and opinionatedly. You've designed multi-step agentic systems and made them reliable, observable, and cost-bounded.You have strong evaluation judgment. You build representative datasets, meaningful baselines, useful error taxonomies, and release criteria that connect offline measurements to production behavior.You treat cost and latency as product constraints. You make measured quality, latency, and cost tradeoffs and use the appropriate combination of model selection, prompting, retrieval, caching, distillation, and fine-tuning rather than reaching reflexively for a more complex model.You operate autonomously on ambiguous problems. Given a rough product idea, a few customer quotes, and a Slack thread, you come back with a plan, a prototype, milestones, and a point of view on tradeoffs.You contribute beyond your domain. As a senior IC, you go into whatever part of the codebase a problem requires, recognize issues beyond your immediate area, and translate between engineering goals and business objectives.You up-level the people around you. You mentor by pairing on hard problems, providing substantive design and code reviews, and spreading agent engineering literacy across the team.You're customer and product-driven. You're comfortable on customer calls and in feedback threads, you turn raw signals into requirements, scopes, and milestones.You're pragmatic, not a perfectionist. You ship the smallest correct thing, prefer robust solutions over complicated ones.On the engineering fundamentals:You're a strong software engineer who can ship production services.You're comfortable across our stack - Go on the backend, TypeScript on the frontend, GraphQL, Postgres, Docker - or you're clearly able and eager to get there.You're fluent with agentic coding tools, and you understand and own every line it submits.You're comfortable in an async-first, multi-service, fast-paced remote environment.Nice-to-haves:You've shipped an LLM-powered or agentic developer-facing product you can speak about opinionatedly.You've fine-tuned, distilled, or trained models to meet cost, latency, or quality targets in production.Experience with retrieval, ranking, embeddings, or search relevance.Experience working directly with enterprise customers and translating their needs into a product.Experience mentoring or up-leveling engineers, especially raising a team's agent engineering fluency.Level📊 This job is an IC4.CompensationWe pay above-market salaries... Your base salary is determined by the IC4 pay band for your location zone (1-4). During the recruiting process, we'll discuss the range applicable to you based on job level, relevant skills, experience, qualifications, and location zone.The starting salary for the IC4 pay band in each zone is:Zone 2: $176,000 USDZone 3: $132,000 USDZone 4: $88,000 USDIn addition to competitive cash compensation, we offer meaningful equity and generous perks and benefits.Interview process We expect the interview process to take 4.75 hours in total.👋 Introduction Stage[30m] Recruiter Screen[45m] Hiring Manager Screen / Resume Deep Dive🧑💻 Team Interview Stage[60m] Technical Interview[60m] Technical Interview[60m] Cross-functional team collaboration / Values🎉 Final Interview Stage [30] LeadershipWe check references and conduct your background check