The coin-flip test: why LLMs fail at calibration, which is the actual product

airesearch12 · x · 2026-09-24

A technical thread arguing most people miss the point about Jev-style agent demos: calibration is the product, and current LLMs fail it. Given a coin toss question, models output skewed distributions like heads(70)/tails(30), and shift to tails(80)/heads(20) when told the previous toss was tails — instead of the correct 50/50 regardless of context. The author asks whether a model can be trained to stay calibrated, and why demos without real decision costs hide this flaw.

Original post →

More from Models

Models channel →