AI Evaluation Engineer at GovWorx | Torre

AI Evaluation Engineer

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Compensation
USD120k - 170k/year
location_on
Remote (for United States residents)
Match
skeleton-gauges
You have opted out of job matches in .
To undo this, go to the 'Skills and Interests' section of your preferences.
Review preferences
Shared by
Emma of Torre.ai
1 day ago

Requirements and responsibilities


About GovWorxGovWorx is helping public safety rise to today's greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-1-1 and emergency communications centers across the country by automating quality assurance, training, and real-time call evaluation—allowing agencies to strengthen their teams and better serve their communities. or the one you already have.Position OverviewWe're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science.You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments.Key ResponsibilitiesDesign, build, and maintain automated AI evaluation pipelines for production LLM applicationsDevelop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methodsBuild offline evaluation datasets and regression testing frameworks to measure AI performance over timeAnalyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvementDesign experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changesDevelop dashboards and reporting that communicate AI quality, reliability, and performance metricsPartner with engineering and product teams to safely deploy and monitor improvements to production AI systemsInvestigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflowsHelp establish best practices for Responsible AI, evaluation methodologies, and continuous model improvementQualificationsMust-HavesMust have US citizenship and pass FBI fingerprint and background check in multiple states3+ years of experience in software engineering, machine learning, data science, or a related technical fieldExperience designing evaluation metrics and interpreting AI model performanceUnderstanding of statistical methods including hypothesis testing and experiment designStrong Python development experienceStrong SQL skills with experience analyzing large datasetsExperience building or supporting production LLM or Generative AI applicationsExperience with prompt engineering and systematic prompt evaluationNice to HaveExperience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label StudioExperience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMakerExperience building dashboards using Tableau, Sisense, Power BI, or similar toolsKnowledge of Responsible AI principles and evaluation methodologiesWhy Join GovWorx?Help build AI systems that directly support first responders and emergency communications professionalsOwn AI quality, evaluation, and continuous improvement for production applicationsWork on cutting-edge LLM technologies and help shape the future of Responsible AICollaborate with a high-performing team across AI, engineering, product, and data scienceSolve technically challenging problems with real-world impact on public safetyInfluence AI strategy and evaluation practices across a growing technology company