Lobehub bench shows high reasoning beats max, blogger slams vendor marketing spin
karminski3 · x · 2026-09-15
Citing Lobehub's internal benchmark (his own tests agree, and even medium beat some quantized small models), karminski3 argues that the vendor marketed max as strongest, then reframed high as the right setting when max underperformed—dressing up weak post-training as "best practices" (use high, Linux bash, patches, watch the CoT). He quips Starbucks wouldn't dare claim a medium has 250ml more than a large.
More from Models
- Streamer tests GPT-6 Astra xHigh on Slay the Spire 2's brutal A10 Ironclad run — Jsevillamol · 2026-09-15
- GPT-6 Built a City Out of Text — Matthew Berman · 2026-09-15
- Reddit User Says Local Model Muse Glimmer Excels at Natural Conversation — New-Pressure-6932 · 2026-09-15
- Astra is 'a beast' at SVG generation, Reddit user says — krzonkalla · 2026-09-15
- Twin prime bound pushed to 186 as GPT-6 Astra launch fuels lab math race — RexDouglass · 2026-09-15
- OpenAI cuts desktop voice pricing ~60%, 2.4x more ChatGPT Voice in Codex — athyuttamre · 2026-09-15