Agent Arena leaderboard adds GPT-5.6 Sol, Terra and Luna near the top
arena · x · 2026-07-28
The Agent Arena leaderboard page shows the new GPT-5.6 Sol / Terra / Luna entries in the top ranks for real-world agentic tasks.
- The leaderboard ranks models by net improvement, confirmed success, steerability, bash recovery, and tool hallucination.
- GPT-5.6 Sol (xHigh) appears near the top with +10.11% net improvement.
- GPT-5.6 Terra (xHigh) and GPT-5.6 Luna (xHigh) also land positive at +4.0% and +3.3%.
- The page currently reflects 1,385,187 sessions across 42 models.
It’s essentially a supporting post for the same benchmark event.
Related event: Agent Arena Adds New GPT-5.6 Variants(2 posts)→
More from Models
- Kimi K3 scales Kimi Linear to 2.8T parameters and drops RoPE for NoPE — rasbt · 2026-07-28
- Llama 405B’s post-human future reads like AI poetry turned up to 11 — aiamblichus · 2026-07-28
- Ben's Bites: Opus 5 Matches Fable 5 at Half Price, Visual Agents in tldraw — Ben's Bites · 2026-07-28
- Kimi K3 reportedly has top-tier KV cache economics among frontier models — teortaxesTex · 2026-07-28
- OpenAI’s Codex now splits into Sol, Terra, and Luna, with Luna priced at one-fifth of Sol — TinfoilTricorn · 2026-07-28
- Alibaba launches the Qwen3.8 Growth Plan after developer feedback on Qwen3.8-Max-Preview — Alibaba_Qwen · 2026-07-28