VulcanBench: Grok 4.5 High Leads, Max Effort Doesn't Equal Better Accuracy

elonmusk · x · 2026-08-22

VulcanBench is a benchmark suite of 100% real engineering tasks that tests across different effort levels. The current leaderboard for Eval Suite 3 shows Grok 4.5 High at the top, followed closely by Fable 5 Low. The author reveals that running models at Max effort often doesn't buy more accuracy but simply costs more and consumes more tokens, taking longer to process.

Original post →

More from Models

Models channel →