AI Applied Scientist (Applied ML & Evaluation Science) at Wizard | Torre

AI Applied Scientist (Applied ML & Evaluation Science)

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Compensation
USD225k - 280k/year
location_on
Remote (for United States residents)
Shared by
Emma of Torre.ai
2 days ago

Responsibilities


About WizardWizard is the top-performing AI Shopping Agent, delivering the best products from across the web with unmatched accuracy, quality, and trust.The RoleWe’re looking for an Applied Scientist to own how we measure, understand, and improve the accuracy of our AI agent. This role sits at the intersection of applied ML, evaluation science, and product. You’ll define what “good” looks like for our agent, build the systems to measure it, and lead the science work to improve it, including fine-tuning the LLM judges that power our evaluation pipeline.You’ll partner with ML Engineering and AI Engineering. What you will do is bring scientific rigor to the most important question at Wizard: is our agent getting better, and how do we know?This is a foundational hire on our science team. Evaluation is the starting point, and the role is scoped to grow into broader applied science work as the surface area of the agent expands (recommendations, personalization, ranking, multimodal, conversational understanding).What You’ll DoDefine and evolve accuracy metrics across the full shopping experience (retrieval, ranking, recommendations, outcomes)Design and run experiments to measure improvements and regressionsBuild and maintain evaluation datasets, benchmarks, and scoring frameworksImprove the LLM judges that power our evaluation pipeline: prompting, calibration, and fine-tuning where it mattersTranslate ambiguous product questions into clear, measurable hypotheses and analysisPartner with ML Engineers to validate model changes and guide iterationIdentify failure modes and edge cases, and drive improvements through dataMake agent performance visible, trusted, and actionable across product and engineeringFirst 3 monthsGo deep on the agent, the current eval pipeline, and the metrics we use todayAudit existing accuracy metrics and benchmarks; identify gaps, blind spots, and signals that aren’t trustworthyBuild relationships with ML, AI Engineering, and ProductShip one quick win: a missing benchmark, an improved metric, or a fix to a misleading signalEstablish a baseline view of agent performance the team can rally aroundMonths 3 to 6Own the evaluation framework: datasets, metrics, scoring, reporting, both offline and onlineDrive measurable improvements to LLM judge quality (calibration, fine-tuning where appropriate)Run experiments that influence at least one significant model or product changeStand up automated evaluation the team trusts before and after every launchBuild dashboards and reporting that make agent performance legible to leadershipBeyond 6 monthsLead applied science work on the next frontier as the agent grows: multi-turn evaluation, multimodal, personalization, ranking quality, conversational understandingInfluence team-level strategy on what we measure, what we improve, and whyMentor and help grow the science function as it expandsWhat Success Looks LikeClear, trusted accuracy metrics are consistently used across product and engineeringA robust automated evaluation framework for both offline and live experimentsModel and product changes are consistently measured before and after launchDemonstrable improvements in LLM judge quality and eval coverageScience leadership that informs what we build, not just whether it worksCareer GrowthDepth track: become the org’s authority on AI evaluation: eval strategy, judge models, agent benchmarkingBreadth track: expand into other applied science problems (recommendations, personalization, ranking, multimodal, conversational understanding) as those areas come onlineLeadership track: Senior / Staff Applied Scientist, with technical leadership across the science functionAs the agent gets more capable, the science problems get richerIdeal Background5+ years in Applied ML, AI Research, or Applied Science (PhD or equivalent depth strongly preferred)Hands-on experience evaluating modern AI/ML systems: LLMs, agents, ranking, or recommendationsDirect experience with LLM-based systems: judge models, RAG, prompt engineering, fine-tuning, RLHF, or similarStrong experimentation foundations: A/B testing, causal inference, statistical rigorProven ability to operate in ambiguity: defining problems, not just solving pre-defined onesClear, structured communication that influences across ML, engineering, and productCompensation & BenefitsThe expected base salary range for this role is $225,000 - $280,000 USD, and will vary based on skills, experience, role level, and geographic location. Final compensation will be determined by considering these factors alongside overall role scope and responsibilities.In addition to base salary, Wizard offers:Equity in the form of stock optionsMedical, dental, and vision coverage401(k) planFlexible PTO and company holidaysFully remote work within the United StatesPeriodic company offsites and team gatheringsWizard is committed to fair, transparent, and competitive compensation practices.Apply directly on RemoteJobs.org: https://remotejobs.org/remote-jobs/ai-applied-scientist-wizard