Jev Does Not Play Dice: 83% confidence, 19% accuracy on a fair die roll

kh-ai · reddit · 2026-09-25

A calibration test exposes Jev's blind spot with known-probability inputs: on 400 fair die rolls it picked face 1 every time at 82.9% mean confidence (19% accuracy), reported 92% on a fair coin (52% right), and compressed a documented 30% shortage risk to 5% via Choice. Key insight: MMLU-style tests check if a model knows a question is hard; the dice test checks if it knows an outcome is unknowable from the input—Jev is weak at the latter. Write-up, code, and data included.

Related event: Dice tests expose LLM calibration failures(2 posts)→

Original post →

More from Models

Models channel →