16-model eval finds Sonnet 5 hedges most, landing in ambiguous zone 41.3% of the time

AlexKim · x · 2026-09-19

A follow-up finding from the same 16-model eval: the metric the author almost skipped turned out most discriminating — how often each model places its answer in the ambiguous 0.35-0.65 middle:

Then a cliff — Opus 5 hedges far less than the rest.

Related event: 16-model calibration test: Jev fastest honest model but ranks 10th in accuracy(8 posts)→

Original post →

More from Models

Models channel →