About AnthropicAnthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.About the roleAs a Staff AI Engineer on the GTM Claudification team, you will build the agents and AI systems that run Anthropic's own go-to-market work. Our sellers already work alongside agents every day. You will take things to the next step and build agents that run complete autonomous motions across areas like inbound, outbound, and pipeline management. In addition, you'll build eval frameworks that prove those agents are ready for customer-facing work and are driving value. This is a senior role where you'll drive technical direction for agents and evals across our team.Working closely with sellers, RevOps, and our platform engineering partners, you'll own projects from first prototype through production operation. You'll combine full-stack engineering (MCP servers, agentic systems, web applications, etc.) with hands-on evaluation work (behavior benchmarks, production monitoring, ROI measurement), and help architect the shared platforms that builders from across our go-to-market org contribute to. You've worked in cultures of analytical rigor before, and you're eager to help shape the norms and best practices of a growing AI engineering function at a pivotal moment in the company's growth.Key responsibilitiesBuild and operate autonomous agents that run go-to-market motions end to end, across areas like inbound, outbound, pipeline management, and customer engagementDesign the human oversight for each motion: approval gates, handoffs, and escalation paths that keep sellers in controlDevelop evaluation frameworks for agent behavior, and run them in development and in productionInstrument model and tool calls in production, and build the observability and measurement that ties agent actions to pipeline and revenueShip MCP servers, agent skills, and web applications that connect to systems like our CRM, communication tools, and data warehouseSet the technical direction for how we build, evaluate, and operate agents across the teamArchitect shared codebases that builders from across go-to-market contribute to, setting the conventions and review practices that keep quality highWork directly with sellers to ground agent designs in real workflows, and iterate based on what you observeIdentify repeatable patterns and contribute insights back to Anthropic's Product and Engineering teamsMaintain strong knowledge of the latest developments in LLM capabilities, agent frameworks, and evaluation techniquesMinimum qualificationsStrong programming skills in Python or TypeScript, with experience building and operating production applicationsProduction experience with LLMs, including context engineering, agent development, MCP development, tool use, and evaluation frameworksExperience using evals and transcript analysis to find and fix real problems in an LLM systemWorking fluency with data, including SQLAbility to navigate ambiguity and ship without a spec, finding simple solutions to complex problemsPassion for advancing safe, beneficial AI, and care for the people who use what you buildPreferred qualifications8+ years in roles such as software engineer, ML engineer, or forward deployed engineer. Former technical founders are encouraged to applyExperience with the Claude Code and the Claude Agent SDKExperience with go-to-market systems (CRM, sales engagement, enrichment, conversation intelligence) or time working closely with a revenue teamExperience growing a codebase that many people contribute to, inner-source or open-sourceApplied ML and experimentation background: A/B testing, propensity models, recommendations, or causal analysisExceptional communication skills to convey technical concepts to non-technical partners with low egoRepresentative projects(Illustrative of the kind of work, not a project list.)Build an agent that takes a routine sales workflow from first signal to a drafted, human-reviewed actionStand up the eval suite for an agent: seed scenarios, scoring rubrics, and regression runs on every changeShip an MCP server that gives sellers and their agents governed access to a core revenue systemDesign a shared repository where go-to-market builders publish agents and skills, with the tests and review rules that keep it healthyBuild a predictive model that explains itself, so an agent can tell a seller why it suggests an action