GPT-6 scores below Kimi K3 on GDPval-AA v2 and neither OpenAI nor Artificial Analysis will explain why

ChrisGPT · x · 2026-09-08

The author puzzles over Artificial Analysis standing by its finding that GPT-6 underperforms Kimi K3 on the real-world GDPval-AA v2 benchmark, while OpenAI stays oddly quiet despite just declaring AGI. Plausible explanations: scoring mistakes, or models answering correctly in ways the benchmarkers didn't account for. Private benchmarks still show GPT-6 clearly ahead on general reasoning, and critics argue OpenAI has benchmaxed elsewhere — but neither side has offered an official explanation.

Original post →

More from Models

Models channel →