GPT-6 Astra tops Mercor's APEX-Agents leaderboard at 62.4% on professional tasks
sherwinwu · x · 2026-09-04
GPT-6 Astra debuts at #1 on Mercor's APEX-Agents leaderboard with a 62.4% mean score and 46.7% Pass@1, edging Claude Fable 5.1 by 0.4 points on mean. It gains 5.7 points on mean and 6.7 on Pass@1 over GPT-5.6 Sol, improving all three professional domains (banking, consulting, law) by 4.5+ points. The benchmark spans 33 simulated work worlds and 480 tasks; the dataset and the Archipelago eval infra are open-source.
More from Models
- OpenAI rolls out misalignment monitoring for Astra, admits it may miss harmful behavior — ShakeelHashim · 2026-09-04
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04