Models Are Getting Dumber on Purpose: Trading World Knowledge for Reasoning
bibryam · x · 2026-08-18
While reasoning scores climb, factual recall is plummeting; Gemini 2.5 Pro leads SimpleQA with just 53%, and small models hallucinate over 80% of the time. Labs are deliberately trading world knowledge for reasoning skills to reduce per-token compute, as facts consume significant parameter space.
More from Models
- Qwen 72B runs on RTX 5090 at 115 tokens/sec, BF16 weights 55.6GB — AccBalanced · 2026-08-18
- Mixedbread releases Toast 1: Beats Fable 5 in deep search, slashes pricing — bclavie · 2026-08-18
- Minimax M3.1 Model Launching Within 48 Hours — ccerrato147 · 2026-08-18
- Rumor: Grok 4.7 to ingest SpaceX engineering data for real-world edge — JOBhakdi · 2026-08-18
- EngramLab model outperforms Opus 4.8 X-high with 3.3x fewer tokens — soumitrashukla9 · 2026-08-18
- Fun observation: Qwen 3.8 thinks like a hardware shopper — Elorun · 2026-08-18