Grok 4.5 Sees Massive Leap in Video Evals
elonmusk · x · 2026-07-13
Elon Musk reshared video studio evaluation feedback regarding Grok 4.5.
The key claims are:
- In the author's video studio evals, Grok 4.5's score jumped from 6/33 to 23/33
- Its cost-efficiency was described as "insane," outperforming the 5.6 series models
This is more of a real-world performance feedback for video workflows rather than an official release note.
Related event: Grok 4.5 First Tests: Faster, Cheaper, and Entering Coding Workflows(8 posts)→
More from Models
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- Qwen 3.8 Next Flash is painfully verbose: 13-minute thinking on single coding prompts — Infinite-Local5435 · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Leaker claims xAI is preparing Grok 4.7, hints at another surprise — mark_k · 2026-09-07
- Local LLMs now near Opus-level — what's still keeping them behind closed models? — mrsalvadordali · 2026-09-07
- Blind test of 12 models finds Fable 5.1 reads least like AI at 14%, Gemini 3.8 Flash worst at 77% — PawelHuryn · 2026-09-07