GPT-6 scores below Kimi K3 on GDPval-AA v2 and neither OpenAI nor Artificial Analysis will explain why
ChrisGPT · x · 2026-09-08
The author puzzles over Artificial Analysis standing by its finding that GPT-6 underperforms Kimi K3 on the real-world GDPval-AA v2 benchmark, while OpenAI stays oddly quiet despite just declaring AGI. Plausible explanations: scoring mistakes, or models answering correctly in ways the benchmarkers didn't account for. Private benchmarks still show GPT-6 clearly ahead on general reasoning, and critics argue OpenAI has benchmaxed elsewhere — but neither side has offered an official explanation.
More from Models
- Astra's home-grown chess engine beats Wally, finds forced mate in 2 by move 29 — MikePFrank · 2026-09-08
- Agent midway to AGI gives up writing full sentences, author jokes about blaming post-training — yacinelearning · 2026-09-08
- Yacine: GPT-6 Astra Gives Up on Unseen Tasks Out of Pure Dread — yacinelearning · 2026-09-08
- Grok Build ships triple daily updates: first-party MCP server, persistent subagents push toward full agent workspace — elonmusk · 2026-09-08
- GPT-6 Astra autonomously rebuilds 20.8km Nürburgring in Blender with 65,000 trees — reach_vb · 2026-09-08
- How GPT-6 Astra's computer use loop powers its viral Blender 3D world generation — iamrobotbear · 2026-09-08