AI Benchmark Engineer and LLM Evaluation Specialist with 1+ years of hands-on experience designing adversarial benchmarks, rubric-based grading systems, and multimodal evaluation pipelines for frontier artificial intelligence models. Proven track record authoring 100+ long-horizon coding tasks with deterministic verifiers and calibrating task difficulty against state-of-the-art large language models including Claude Sonnet and Opus. Expert in prompt engineering, red teaming, content moderation, trust and safety, and quality assurance at scale. Led a team of 15 annotation specialists while maintaining top-tier quality assurance scores for production reinforcement learning from human feedback (RLHF) datasets. Native Arabic speaker with fluent English (C1) and a law degree trained in precise document analysis, policy interpretation, and extracting decisions from ambiguous text. Technical background in Python, SQL, JavaScript, Docker, and bash scripting. Seeking to leverage deep expertise in model evaluation, benchmark infrastructure, and AI alignment to drive performance and safety in next-generation language models.