APEX-Agents 1.1 benchmark update: Claude Fable 5.1 tops leaderboard at 68.6%
amaarora · x · 2026-09-09
- Mercor released APEX-Agents 1.1, refreshing task specs, tooling, and environments to keep the leaderboard accurate as models advance.
- Key change: the benchmark no longer rewards noncommittal answers, aligning grading with how professionals actually operate.
- Pass@1 rankings: Claude Fable 5.1 leads at 68.6%, followed by Gemini 3.7 Flash (67.8%), Claude Opus 5 (65.8%), Grok 4.6 (65.3%), and GPT-6 Astra (64.7%).
- Ryan Marten praised the team's use of GEPA to optimize judges (agentic verifiers need benchmarks too); the benchmark is now on Harbor Hub.
More from Models
- GPT-6 Astra system card draws fire: OpenAI claims 'most aligned model' ever — TheZvi · 2026-09-09
- Muse reportedly makes phone calls in Croatian, but users say it denies the ability — nathanbenaich · 2026-09-09
- DeepSeek V4.1 Flash tops Chinese models in coding blind test, but drops 17 points when switching clients — teortaxesTex · 2026-09-09
- GPT-6 Astra gears up for Singapore F1 night race with live demo — gabrielchua · 2026-09-09
- Nex-N2.5 Goes Free on OpenRouter: 262K-Context Agentic Coding Model with Visual Feedback Loop — airesearch12 · 2026-09-09
- Raschka deep dive: GPT-6 Astra, recurrent depth, and whether it hides its chain of thought — bibryam · 2026-09-09