16-model eval finds Sonnet 5 hedges most, landing in ambiguous zone 41.3% of the time
AlexKim · x · 2026-09-19
A follow-up finding from the same 16-model eval: the metric the author almost skipped turned out most discriminating — how often each model places its answer in the ambiguous 0.35-0.65 middle:
- Sonnet 5: 41.3%
- Fable 5.1: 36.7%
- Jev: 34.7%
- Opus 5: 23.3%
Then a cliff — Opus 5 hedges far less than the rest.
More from Models
- Google finally has a SOTA rogue model, and the AI crowd is joking about relief — rao2z · 2026-09-19
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19
- Empirical analysis confirms Claude Opus 5 shows abnormally dark base-model completions — Kyrannio · 2026-09-19
- US products quietly build on Chinese open-weight models as one firm cuts spend by ~100x — generativist · 2026-09-19
- Eval model Jev goes free on Vercel AI Gateway until Sept 25 — cramforce · 2026-09-19