I build AI systems that make judgement calls in production, and I have run the P&L those systems serve.
Ad Engine is a compliance engine that scores health advertisers against a 10 category ruleset. It ingests the Meta Ad Library, runs every creative through an LLM judge, and cites each finding to an Ad Library ID so anyone can check it at source. I have run it across 14 live US advertisers, scores 0 to 62. PRISM is an AWS pipeline (EventBridge, Lambda, Aurora, Bedrock) collecting campaign creative continuously. SIGNAL flags divergence between reported and real conversions across Meta CAPI, GHL and HubSpot.
I am currently building the evaluation layer: a 528 ad corpus with full denominators, an LLM-as-judge scored against human labels, and a bias corrected violation rate rather than an accuracy number. Python, Next.js, Node, AWS.
Before the tooling I did the work by hand. I grew a hospital partnership from zero to about $215K in partner attributed revenue at a $100 CPA, sourced a pharmacy partner and placed around 50 clinics on placement fee and revenue share terms, and ran creator rosters on commission judged against 90 day LTV to CAC. That is why the systems measure what they measure.
Audits are public at runadengine.com/a