Evaluator
Perle Global Social Media Platform Project NDA
Jan 2024 - Jan 2025 (1 year 1 month)
Evaluated AI-generated Arabic and multilingual responses for tone, safety, relevance, and policy alignment. Reviewed complex multi-step AI agent outputs and reasoning traces for logical consistency and task completion. Identified recurring model errors, hallucination risks, unclear responses, and instruction-following failures. Applied structured rubrics to score response quality and provide actionable feedback for model improvement. Worked under strict confidentiality requirements while handling sensitive AI evaluation workflows.