HF researcher disputes GLiDE's Decision Index win over Jev as reasoning-boosted
antoine_chaffin · x · 2026-10-06
Hugging Face's Antoine Chaffin pushes back on GLiDE topping the Decision Index (64.81 vs Jev's 57.91 over 155k+ questions): GLiDE likely re-enables generative reasoning, which inflates the knowledge/reasoning sections but deviates from the models' original purpose. He checked that keeping backbone perf on HLE etc. is easy if you just turn reasoning on — so the comparison isn't apples-to-apples, though calibrated reasoning activation remains a promising path.
More from Models
- Reflection's Billion-Dollar US Open-Weight Model Underperforms Every Major Chinese Model, Mocked Online — npinto · 2026-10-06
- Opus 5.5 uses fewer tokens than GPT-6.1 Sol while scoring better, fan argues — Angaisb_ · 2026-10-06
- Attention Relay makes embedding models instruction-aware without training via LLM attention weights — _reachsumit · 2026-10-06
- OpenWork, an open-source Claude Cowork alternative, hits 45K downloads in two days — alex_verem · 2026-10-06
- Abacus.AI CEO: Chinese open-source models beat US labs on easy tasks, DeepSeek cheapest for agents — bindureddy · 2026-10-06
- ChatGPT quietly adds lifetime usage tracking to profile settings — Bpelks · 2026-10-06