Calibration Metrics Added to Open Leaderboard; Decider 35B-A3B Takes the Lead
multimodalart · x · 2026-09-22
multimodalart added calibration metrics to their open leaderboard, with Decider 35B-A3B (by notmapika), built on Qwen3.5-35B-A3B-Base, taking the lead as the most calibrated model.
- Calibration measures how well a model's confidence matches its actual accuracy — an often overlooked but critical capability
- The entire methodology is open source and reproducible; docs and code are available on GitHub
More from Models
- GPT-6 Astra tops Claude Opus 5 in Ramp business spending share, OpenAI overtakes Anthropic — firstadopter · 2026-09-22
- NYT DealBook confirms OpenAI overtook Anthropic in business AI spending last week — firstadopter · 2026-09-22
- 86 physics questions, 100 runs each: model hits 87.6% accuracy but understates confidence — Ok-Challenge-7810 · 2026-09-22
- DeepSeek reportedly bets on Huawei chips to train next-gen models; Liang says it 'has to work' — kimmonismus · 2026-09-22
- Xiaomi's MiMo-V2.6-Pro tops open models on $2.62M RL; Anthropic alleges Claude distillation — The Decoder · 2026-09-22
- Tencent finally opens WeChat interface, unlocking 100GB+ chat data processing — Xianbao_QIAN · 2026-09-22