Selected Project Experience
Selected Project Experience
Jan 2025 - Current (1 year 8 months)
Terminal-Bench 2.0 — AI Agent Benchmark Tasks
• Built terminal-based development tasks requiring agents to interpret specifications, implement code, run tests, and iterate from feedback.
• Designed validators for oracle correctness, coverage, complexity, and long-horizon behavior.
• Example domains: Linux hardening, build systems, content-addressable caching, sparse matrices, MVCC storage engines, LSM-tree storage, document ranking, and TypeScript CLIs.
• Example validation result: Oracle, Coverage, Complexity, and Long-horizon on a Linux hardening benchmark task.
IaC Audit V3 — Infrastructure-as-Code Review
• Reviewed Infrastructure-as-Code tasks and solutions for correctness, execution quality, test coverage, and evi