Priciest model finishes last on accuracy and calibration in 16-model test
AlexKim · x · 2026-09-19
Price and vintage do not predict calibration: kimi-k3 was the most expensive model measured and finished last on both accuracy and calibration, while deepseek-v4-flash ran the whole set for $0.003 and beat Jev on accuracy. The takeaway: neither price nor pedigree predicts whether a model will admit it is unsure.
More from Models
- Redditor tests typesafeAI's new classifier Jev: roughly GPT-1-level smart, with a trick to gauge its knowledge cutoff — MrCyclopede · 2026-09-19
- RSI-Exam leaderboard update: GPT-6-astra tops recursive self-improvement benchmark at 0.5126 — yuyinzhou_cs · 2026-09-19
- Claude Max plan hit with usage complaints: quota exhausted after just four hours — yangyi · 2026-09-19
- Grok Bot voice mode now fully rolled out on mobile — Daniel_Farinax · 2026-09-19
- AfriqueQwen 3.5 post-trained models released, claiming wins over Gemma 3 and Apertus — davlanade · 2026-09-19
- Karpathy's year-old rant calling AI agents "slop" resurfaces as agent hype rolls on — steipete · 2026-09-19