Grok 4.7 lands in Agent Arena; community votes on millions of real agentic tasks
arena · x · 2026-09-22
Grok 4.7 is now in LM Arena's Agent Arena, which measures models on millions of real-world, long-horizon agentic tasks voted on by a global community. Models can use web search, filesystem and terminal tools to complete complex workflows, and the leaderboard scores outcome performance relative to the average model using causal tracing. Grok 4.7 has also entered Battle Mode for Text, Vision, Code and Document, with scores coming soon as votes come in.
More from Models
- Building an image rating tool with GPT Vision and Jev: what worked and what didn't — huangyun_122 · 2026-09-22
- Researcher: Gemini's 'depressive spirals' during tasks deserve systematic investigation — jachiam0 · 2026-09-22
- Harrison Chase on decision models: agentic systems are just good engineering around models — Hacubu · 2026-09-22
- ChessLFM makes it to the Lichess home page, closing the loop on its seed data source — maximelabonne · 2026-09-22
- A 27B model one-shot recreates 618 visual styles as SVG artworks, text-only — Seromelhor · 2026-09-22
- Grok 4.7 now usable in Cursor, early hands-on — IndraVahan · 2026-09-22