TRACES Board: GPT-6-astra Tops Agent Leaderboard but Trails Opus-5 on Self-Repair
thisdudelikesAI · x · 2026-09-12
The TRACES leaderboard launched a six-capability agent evaluation (0-4 scale: tools, repair, alternatives, coherence, evidence, scope), running 140 agent episodes under a shared harness.
Key points:
- GPT-6-astra tops the board, beating GPT-5.6-sol across all six capabilities, but unevenly: tools 2.34→2.80, alternatives 2.10→2.80, coherence 2.63→3.15, evidence 2.26→2.76, while repair only moved 1.81→2.01.
- Category leaders differ: GPT-6-astra leads tools (2.80), Opus-5 leads repair (2.59) and alternatives (2.85), with GLM-5.2 and Kimi-k3 close behind.
- Kimi-k3, GLM-5.2, and DeepSeek-v4-flash all place prominently, highlighting structural capability differences among frontier models.
More from coding & agent
- Stripe ships Link CLI: add agentic payments to your app with a one-shot integration — jeff_weinstein · 2026-09-13
- Dev feeds GPT-6 Gran Turismo and F1 movie to build an interactive 3D racing scene — cedric_chee · 2026-09-13
- Will coding agents invent their own formal languages, with natural language as the metalanguage? — jmugan · 2026-09-13
- Uncle Bob Publishes 'Morning Bathrobe Rant' Rethinking AI Coding Harnesses — blaizedsouza · 2026-09-13
- Poison Event Quarantine: 6-Step Cheatsheet to Stop Bad Payloads Blocking Your Queue — blaizedsouza · 2026-09-12
- Fanout, Not Direct Calls: One Write, Many Readers for Production Agents — blaizedsouza · 2026-09-12