Grok 4.7 lands in Agent Arena; community votes on millions of real agentic tasks

arena · x · 2026-09-22

Grok 4.7 is now in LM Arena's Agent Arena, which measures models on millions of real-world, long-horizon agentic tasks voted on by a global community. Models can use web search, filesystem and terminal tools to complete complex workflows, and the leaderboard scores outcome performance relative to the average model using causal tracing. Grok 4.7 has also entered Battle Mode for Text, Vision, Code and Document, with scores coming soon as votes come in.

Original post →

More from Models

Models channel →