Not every model failure is lack of capability: Terminal-Bench evals hide safety declines

abeirami · x · 2026-09-30

When evaluating models on Terminal-Bench, researcher ijulin argues that accuracy alone hides an important distinction: not every failure stems from lack of capability — some reflect safety-driven performance declines. The thread (with a comparison figure) calls for attributing failures to capability vs. safety separately in evals.

Original post →

More from Models

Models channel →