Mercor's APEX-1: GPT-5.6 Terra tops 400-task professional benchmark
zainhas · x · 2026-09-09
Mercor released and open-sourced APEX-1, a benchmark testing frontier models on economically valuable work across four professions—investment banking, consulting, big law, and primary care—with tasks written by veteran practitioners and graded 8x per task by a Judge LM on 400 hidden tasks.
- Top 5: GPT-5.6 TerraMax 69.5% ±2.4%, GPT-6 AstraXHigh 69.3%, Opus 5 XHigh 68.6%, Kimi K3 Max 68.4%, GPT 5.4 High 67.2%—all within error bars
- Legal track advised by Cass Sunstein, sourced from Latham & Watkins, Skadden, Cravath
- 100 in-distribution cases, eval harness, and paper are open-sourced on Hugging Face
zainhas flagged odd rankings: GPT 5.6 Terra above Astra, and GPT 5.4 beating Fable 5 while tying Gemini 3.6 Flash.
Related event: Mercor's APEX-1 Benchmark Released, Rankings Draw Skepticism(2 posts)→
More from Models
- Nex N2.5 mini, a Qwen3.5 MoE-based image-text-to-text model, trends on Hugging Face — nex-agi · 2026-09-09
- Nex-N2.5-Pro released under Apache-2.0 and trends on Hugging Face — nex-agi · 2026-09-09
- Nvidia's Jensen Huang: closed models are cheaper, open models give you control — rohanpaul_ai · 2026-09-09
- Analyst speculates OpenAI achieved multi-agent swarm breakthrough, with agents self-organizing since May — teortaxesTex · 2026-09-09
- GPT-6 called the most interesting model release in years, thanks to HF and German forum leaks — xeophon · 2026-09-09
- Anthropic Max users report Opus quietly excluded from 'all models' usage bar — wyongriver · 2026-09-09