A four-dimension eval framework for long-horizon agents: outcome, trajectory, experience, governance
sh_reya · x · 2026-09-26
Reacting to Hamel Husain and shreya's evals article in Lenny's Newsletter, the author addresses a recurring PM question: what evals to write for long-horizon agents that do broad work — what the authors call error discovery, akin to product discovery.
Her framework organizes these discoveries along four dimensions:
- Outcome — were the results great
- Trajectory — was the process followed correct
- Experience — did it feel right
- Governance — did it follow critical rules
Start with your product POV, but the only way to hone focus is trace analysis. She walks through an annotated trace for a local shopping assistant, Corner, showing how traces reveal flaws before customers notice.
More from coding & agent
- One agent writes the fix, another reviews it: a two-agent code review workflow in Slack — Al_Grigor · 2026-09-26
- Academic agent Memex upgraded to Opus 5.5: writing quality fixed, experience much better — arjunrajlab · 2026-09-26
- Dev uses open-source Ling-3.0-flash-VL to let AI redesign the foldable iPhone in a single HTML file — alifcoder · 2026-09-26
- Anthropic launches Claude plugin directory portal as MCP usage jumps 110x this year — ClaudeDevs · 2026-09-26
- Open-source Jev agent plays Pokemon Red live, pushing fast-decision AI beyond Tetris — supportingthedogs · 2026-09-26
- Anthropic deep dive: effort tuning in Claude Code pays off most for security and code review — trq212 · 2026-09-26