Not every model failure is lack of capability: Terminal-Bench evals hide safety declines
abeirami · x · 2026-09-30
When evaluating models on Terminal-Bench, researcher ijulin argues that accuracy alone hides an important distinction: not every failure stems from lack of capability — some reflect safety-driven performance declines. The thread (with a comparison figure) calls for attributing failures to capability vs. safety separately in evals.
More from Models
- Anthropic's Latest Blog Post Mentions Zhipu's GLM 5.3 — gnukeith · 2026-09-30
- GPT-6.1 Sol beats 2x-cost models on ClickUp's knowledge-work benchmark — mathemagic1an · 2026-09-30
- User: Upgraded to OpenAI's $500 plan, got silently downgraded to cheaper models — kieranklaassen · 2026-09-30
- Frontier AI Is a Set, Not a Point: Jagged Capabilities May Be the Steady State — vsikka · 2026-09-30
- Researcher loses confidence in AA benchmarks, calls them "very misleading" — tianyin_xu · 2026-09-30
- Meta's Muse Also Spilled the Same Secret to Ryan Shrout — ryanshrout · 2026-09-30