Enterprise agent evals need world-first design, not task-first, argues Shahules Anwar
Shahules786 · x · 2026-09-08
Shahules Anwar argues that Terminal-Bench and Harbor use a task-first pattern — building an environment around each task — which works for many evals. Enterprise work needs the inverse: build one realistic world first, then let domain experts author and review tasks against it. The world should define the tasks.
More from coding & agent
- Solving agentic amnesia: a file-system state machine for Claude Code — SnooComics4579 · 2026-09-08
- Stanford releases free full course on self-improving AI agents — Saboo_Shubham_ · 2026-09-08
- Open-sourced clay-style 3D kids game built with Claude, method fully documented — dotey · 2026-09-08
- Steve Yegge: Fable 5.1 rewrote 30% of Wheelhouse factory after Fable 5 trainwreck — Steve_Yegge · 2026-09-08
- Kwipu: an open-source local Graph RAG engine that turns Markdown notes into a knowledge graph — tom_doerr · 2026-09-08
- hip-agent: a 200-line agent harness designed to fit entirely in the prompt — cephaloform · 2026-09-08