The coin-flip test: why LLMs fail at calibration, which is the actual product
airesearch12 · x · 2026-09-24
A technical thread arguing most people miss the point about Jev-style agent demos: calibration is the product, and current LLMs fail it. Given a coin toss question, models output skewed distributions like heads(70)/tails(30), and shift to tails(80)/heads(20) when told the previous toss was tails — instead of the correct 50/50 regardless of context. The author asks whether a model can be trained to stay calibrated, and why demos without real decision costs hide this flaw.
More from Models
- User claims MiniMax H3 was trained on The Will Stancil Show after model adds unprompted whistle — Kyrannio · 2026-09-24
- Veteran Demoscene Dev Defends Opus 5.5 One-Shot 90s-Style Demo Amid Skepticism — gandamu_ml · 2026-09-24
- Opus 5.5 One-Shots a 90s-Style Demoscene Demo in C/C++ and OpenGL — BugoTheCat · 2026-09-24
- Xiaomi's MiMo V2.6 Pro builds a habit tracker in 62 seconds for under 5 cents — socialwithaayan · 2026-09-24
- Gemini 4 Is Almost Ready, Says New Google DeepMind Chief Koray Kavukcuoglu — The Verge AI · 2026-09-24
- New chat, no refusal: user shows ChatGPT's self-image prompt only blocked in original thread — Todesluke · 2026-09-24