GPT-6 Regresses on GDPval Knowledge-Work Benchmark Due to Shorter Deliverables

ArtificialAnlys · x · 2026-09-23

Follow-up from Artificial Analysis: both GPT-6 Sol and Luna regress on GDPval-AA v2.1 at max effort — a benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations. Manual inspection attributes regressions to shorter deliverables that more often omit required rubric elements.

Related event: GPT-6 Sol divides reviewers: gains in coding agents, regression on DeepSWE, at half the price(11 posts)→

Original post →

More from Models

Models channel →