VulcanBench: Grok 4.5 High Leads, Max Effort Doesn't Equal Better Accuracy
elonmusk · x · 2026-08-22
VulcanBench is a benchmark suite of 100% real engineering tasks that tests across different effort levels. The current leaderboard for Eval Suite 3 shows Grok 4.5 High at the top, followed closely by Fable 5 Low. The author reveals that running models at Max effort often doesn't buy more accuracy but simply costs more and consumes more tokens, taking longer to process.
More from Models
- GLM 5.3, Fable 5, and GPT-5.6 Sol show opposite results on Terminal-Bench 3 vs DeepSWE — zainhas · 2026-08-22
- Claude interrogates you to guess your vibe; Grok just reads your tweets — repligate · 2026-08-22
- Relying solely on benchmarks and consensus fails to capture true model capabilities — nptacek · 2026-08-22
- Opus 5 allocates skills to coding, philosophy, and understanding human intent — davidad · 2026-08-22
- Fable 5 excels at postdoc-level math, reversing Anthropic's historical underperformance — davidad · 2026-08-22
- Frontier model capabilities are jagged; custom evals for specific use cases are essential — nptacek · 2026-08-22