Mistral Large 4 burns over 2x the output tokens per task vs GPT-6 sol

haider1 · x · 2026-10-06

Independent evaluator haider reports that Mistral Large 4 uses over 2x as many output tokens per Intelligence Index task as GPT-6 sol and Astra — costly for a model that isn't leading on intelligence. His earlier comparison placed Mistral Large 4 on par with leading open-weight models, likely the strongest Western open-weight model, though Mistral hasn't published a full benchmark comparison.

Original post →

More from Models

Models channel →