Jev Does Not Play Dice: 83% confidence, 19% accuracy on a fair die roll
kh-ai · reddit · 2026-09-25
A calibration test exposes Jev's blind spot with known-probability inputs: on 400 fair die rolls it picked face 1 every time at 82.9% mean confidence (19% accuracy), reported 92% on a fair coin (52% right), and compressed a documented 30% shortage risk to 5% via Choice. Key insight: MMLU-style tests check if a model knows a question is hard; the dice test checks if it knows an outcome is unknowable from the input—Jev is weak at the latter. Write-up, code, and data included.
Related event: Dice tests expose LLM calibration failures(2 posts)→
More from Models
- Vercel gateway data: Anthropic spend share falls 69%→40% as OpenAI doubles — ctjlewis · 2026-09-25
- Muse Spark 1.3 Now Available via Oracle, in Private Preview on Google Cloud — alexandr_wang · 2026-09-25
- Gemini 3.8 Flash scores 89.2% on ARC-AGI-2 at $0.40/task, 98.5% on ARC-AGI-1 — fchollet · 2026-09-25
- Arena Launches Redesigned Leaderboard Hub Unifying Live Model Eval Signals — arena · 2026-09-25
- Early verdict: Opus 5.5 is the best Claude yet — well-rounded and faster, weaker at math — marcosalvi · 2026-09-25
- Using Jev-style system-1 models as a cheap calibrated decision layer for Bittensor validators — markjeffrey · 2026-09-25