LLM fails coin-flip probability test, says 60/40 coin is 99/1
JnBrymn · x · 2026-09-19
Developer JnBrymn stress-tested a model called Jev on probability calibration: for a coin that lands heads 60% of the time, the model claimed 99/1. Scanning all possible head probabilities showed broadly poor calibration. Reframing the question as a 'noul' rather than a 'choice' made outputs saner, but max error on questions with known exact answers was still 9% — highlighting systematic miscalibration and prompt sensitivity.
Related event: Jev model flunks probability calibration test, says 60/40 coin is 99/1(2 posts)→
More from Models
- Dev finds DeepSeek 4.1 Flash faster and cheaper than Gemini Flash — but drops it over data concerns — julianharris · 2026-09-19
- RL-trained classification model Laya trends on Hugging Face — convaiinnovations · 2026-09-19
- OpenAI adds new voices to GPT-Live API, available in developer playground — juberti · 2026-09-19
- DeepSeek v4.1 flash paper figure shows sharp quality jump at 1M-token context — andrew_n_carr · 2026-09-19
- User charged $1,600 for idle containers on OpenAI's new Agents API — Wide-Arugula3042 · 2026-09-19
- Benchmarks show vision models spot image patterns but fail to pick the correct next image — AndrewDai · 2026-09-19