AI Agent Evaluation Specialist at Outlier AI | Torre

AI Agent Evaluation Specialist

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Freelance
Recurrent
Compensation
USD2.4k - 5k/month
Negotiable
location_on
Remote (anywhere)
Posted 14 days ago

Responsibilities


As an AI Agent Training Contributor, you may work on tasks involving: - Evaluating AI-generated responses and agent behavior. - Reviewing complex, multi-step AI workflows. - Identifying logical, factual, technical, or execution errors. - Testing AI agents against edge cases and failure scenarios. - Analyzing agent traces, tool usage, and reasoning paths. - Creating and applying evaluation criteria and rubrics. - Providing structured feedback to improve AI accuracy and reliability. - Reasoning through software, data, QA, and AI-related problems. - Helping improve the performance of next-generation autonomous AI agents. Key Responsibilities: - Review AI-agent outputs for correctness, relevance, consistency, and instruction-following. - Detect hallucinations, reasoning failures, workflow errors, and incorrect tool usage. - Evaluate whether an AI agent successfully completes multi-step technical tasks. - Compare alternative AI responses and determine which performs better. - Test models using realistic scenarios, adversarial prompts, and edge cases. - Document issues clearly and explain why a response or workflow succeeds or fails. - Follow detailed project guidelines, rubrics, and evaluation standards. - Provide high-quality written feedback that can be used to improve AI models. - Perform technical analysis involving code, software workflows, data, automation, or AI systems when required. - Maintain consistency, accuracy, and attention to detail across evaluation tasks. - Adapt to different task types and project-specific instructions. - Complete qualification assessments and maintain required quality standards.