Harrison Chase shares strategies to harden Agent evals for better signals
hwchase17 · x · 2026-08-21
Harrison Chase shared an update on Eval Engineering, discussing how to make eval tasks more realistic and challenging. Strategies include adding search-requiring information discovery and multi-case completeness checks rather than just increasing data volume. Comparing weak and strong model performances helps calibrate task difficulty, providing better learning signals for Agent optimization.
More from coding & agent
- AI coding agents cannot replace senior engineering thinking yet — bendee983 · 2026-08-22
- Moving to Cloud Agents: Solving Local Conflicts and Dependency Pain Points — jarrodwatts · 2026-08-22
- NEURA Robotics runs daily evals across ~100 robot cells to teach robots to know when they're stuck — wandb · 2026-08-22
- Ox Alpha scores 96% on SWE-bench, author skeptical — No_Tip9917 · 2026-08-22
- ThreeUI open-sources 160+ three.js components with prompts to feed your agent — dankaplan · 2026-08-22
- LangChain.js beginner course updated with Microsoft Foundry, works with any OpenAI-compatible API — DanWahlin · 2026-08-22