xAI Launches Grok Imagine Video 1.5 Agent, Ranking 5th on Text-to-Video Arena
On September 6, xAI officially announced that Grok Imagine Video 1.5 agent is now live. The new version is powered by the latest Image 2.0 model, delivering higher video quality, a smarter agent, and stronger storytelling ability, with particular strength in stitching multiple shots together more coherently. Around the same time, the model entered LMArena's Text-to-Video Arena leaderboard, ranking 5th with a score of 1491, on par with Wan-3.
Confirmed
- Grok Imagine Video 1.5 agent was officially launched by xAI, powered by the Image 2.0 model
- Three improvements highlighted by the official announcement: better visual quality, a smarter agent, and enhanced multi-shot coherence and narrative capability
- Ranked 5th on the Text-to-Video Arena leaderboard with a score of 1491
- xAI research lead Jeff H. reshared and thanked Yunzhi Zhang for leading the agent intelligence work; the release was completed by his team
Why it matters
- Multi-shot coherence has long been a weak spot in AI video generation; making it the core selling point of this release shows xAI's targeted push into narrative-driven video generation
- Entering Text-to-Video Arena's top five puts it in direct competition with leading models like Wan-3
- Users (e.g., real-world feedback relayed by DanielFarinax) report a perceptible improvement in longer-form video generation, suggesting the gains go beyond official marketing claims
2026-09-06 ~ 2026-09-06 · 6 related posts
Primary sources
- [source] xAI ships Grok Imagine Video 1.5 agent with better multi-shot continuity — Daniel_Farinax · 2026-09-06
- xAI teases Grok Imagine Video 1.5 agent led by Yunzhi Zhang — ZhitingHu · 2026-09-06
- [source] Grok Imagine Video 1.5 agent lands #5 on Text-to-Video Arena with 1491 pts, 3 pts behind leaders — jefffhj · 2026-09-06
- Grok launches Imagine Video 1.5 agent powered by Image 2.0 for multi-shot storytelling — belce_dogru · 2026-09-06