Mistral Large 4 burns over 2x the output tokens per task vs GPT-6 sol
haider1 · x · 2026-10-06
Independent evaluator haider reports that Mistral Large 4 uses over 2x as many output tokens per Intelligence Index task as GPT-6 sol and Astra — costly for a model that isn't leading on intelligence. His earlier comparison placed Mistral Large 4 on par with leading open-weight models, likely the strongest Western open-weight model, though Mistral hasn't published a full benchmark comparison.
More from Models
- Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason — arankomatsuzaki · 2026-10-06
- Best move for Anthropic? Stay silent and drop 'Fable 5.5' while OpenAI stumbles — weswinder · 2026-10-06
- New model release mocked as set to be beaten by Qwen 4 27B at 8x smaller size — gnukeith · 2026-10-06
- M5 Mac 128GB local LLM benchmark: Qwen3.8-flash-next hits 40-60 tok/s, Splash hits 120 — surrealerthansurreal · 2026-10-06
- Mistral releases Large 4, a 1T-parameter multimodal model aiming to leapfrog rivals — TechCrunch AI · 2026-10-06
- Heavy Claude User on Gemini 4 Argon: Doesn't Beat Claude for Code, Antigravity Feels Alien — MicahBerkley · 2026-10-06