Why does GPT 5.6 Terra beat Astra on the APEX-1 leaderboard?
zainhas · x · 2026-09-09
zainhas questions Mercor's APEX-1 leaderboard, where GPT 5.6 Terra ranks above Astra, and GPT 5.4 outscores Fable 5 while tying Gemini 3.6 Flash—results that clash with common expectations. APEX-1 tests models on realistic tasks from four professions: investment banking, consulting, big law, and primary care.
Related event: Mercor's APEX-1 Benchmark Released, Rankings Draw Skepticism(2 posts)→
More from Models
- Returning Users Ask: Are ChatGPT Astra's Lowest-Tier Limits Actually Usable? — DragonsRageRS · 2026-09-09
- Anthropic Limits Claude Mythos 5.1 to US Orgs, Cutting Out UK's AISI — Miles_Brundage · 2026-09-09
- Polymarket bets ~93% that OpenAI will preemptively reset Codex weekly limits — Polymarket · 2026-09-09
- OpenAI's 80% price cut drove 10-13x usage, squeezing Anthropic's premium IPO story — rohanpaul_ai · 2026-09-09
- Tencent Hunyuan Hy4 preview hailed as frontier-level, free on WorkBuddy — TencentHunyuan · 2026-09-09
- Schmidhuber: long-context autoregression was cool long before 2016, cites pioneer-credit survey — SchmidhuberAI · 2026-09-09