Grok 4.5 Shows Strong Software Benchmark Performance
BWay124 · x · 2026-07-13
This post relays/responds to model info regarding Grok 4.5, noting that it is "even slightly higher than Fable" on certain software benchmarks. The author adds that they have used it exclusively for days, despite being a long-time user of Claude Pro Max and Cursor.
While the original post is highly anecdotal, the core focus remains on the model's comparative performance on software benchmarks and the fact that it has entered the stage of real-world usage switching.
Related event: Musk Says Grok 4.5 Beats Fable on Some Coding Benchmarks(3 posts)→
More from Models
- A 2-minute Astra audit at low setting wiped a Plus user's full 5-hour limit — Existing-Slide7395 · 2026-09-07
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- Qwen 3.8 Next Flash is painfully verbose: 13-minute thinking on single coding prompts — Infinite-Local5435 · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Leaker claims xAI is preparing Grok 4.7, hints at another surprise — mark_k · 2026-09-07
- Local LLMs now near Opus-level — what's still keeping them behind closed models? — mrsalvadordali · 2026-09-07