Deep dive: 'Evals Are Bullsh*t' - why AI evaluations miss the point
hamostaf04 · x · 2026-08-02
X user hamostaf04 shares an article titled 'Evals Are Bullsht'. The article argues that most AI evals grade answers, but businesses live with decisions. Evals seem objective but often fail to answer the key question: is the agent doing the job the business actually needs? Graders may produce consistent scores, but they don't reflect how your company works, which mistakes are expensive, or what your best operator would have done. The article calls for rethinking evaluation.
Related event: Rethinking AI Agent Evals: Business Decisions Matter More Than Text(4 posts)→
More from AGI Musings
- Logan: You're Probably Underestimating the Exponential Slope of AI Models — OfficialLoganK · 2026-08-03
- AI Risk Forecasts from 34 Sources Show Worsening Trends Year Over Year — avoidthe9to5 · 2026-08-03
- Survey: 65% of Workers Miss the Pre-AI Workplace, 38% Want to Erase GenAI — VraserX · 2026-08-03
- AXRP Podcast Features Eli Lifland Discussing AI 2027 Forecasts — dfrsrchtwts · 2026-08-03
- DARPA's 1983 AI Strategy Plan Included Autonomous Vehicles and AI Copilots — frankreddit5 · 2026-08-03
- Next Telecom Revolution: 5G/6G Networks to Perceive Without Cameras — mustafamhus · 2026-08-03