Summary AI Evaluation Specialist with 3+ years of experience evaluating AI systems across text, image, audio, video, search, and agentic workflows, with 17,500+ AI evaluations completed across multiple AI platforms. Experienced in comparative evaluation, detailed guideline interpretation, quality review, edge-case analysis, and testing AI agents for reasoning, tool use, workflow execution, safety, and instruction following.
Skills & Evaluation: Multimodal AI Evaluation • Comparative Model Evaluation • Preference Ranking & Human Feedback • Rubric Development • Quality Review & Reviewer • Calibration • Guideline Interpretation • Reasoning Evaluation • Hallucination Detection • Edge-Case Analysis • Red Teaming & Adversarial Testing • Model Behavior Assessment • Browser-Based AI Agent Testing • Workflow Simulation • Task & Scenario Design • Tool Use Evaluation
Technical Tools: SQL • JSON • HTML/CSS • JavaScript • Python (Basic) • Audacity • GIMP • Google Workspace • Microsoft Office • Slack