Agent Arena leaderboard adds GPT-5.6 Sol, Terra and Luna near the top
arena · x · 2026-07-28
The Agent Arena leaderboard page shows the new GPT-5.6 Sol / Terra / Luna entries in the top ranks for real-world agentic tasks.
- The leaderboard ranks models by net improvement, confirmed success, steerability, bash recovery, and tool hallucination.
- GPT-5.6 Sol (xHigh) appears near the top with +10.11% net improvement.
- GPT-5.6 Terra (xHigh) and GPT-5.6 Luna (xHigh) also land positive at +4.0% and +3.3%.
- The page currently reflects 1,385,187 sessions across 42 models.
It’s essentially a supporting post for the same benchmark event.
Related event: Agent Arena: Claude Opus 5 Outperforms GPT-5.6 Variants(4 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23