GPT-6 Regresses on GDPval Knowledge-Work Benchmark Due to Shorter Deliverables
ArtificialAnlys · x · 2026-09-23
Follow-up from Artificial Analysis: both GPT-6 Sol and Luna regress on GDPval-AA v2.1 at max effort — a benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations. Manual inspection attributes regressions to shorter deliverables that more often omit required rubric elements.
More from Models
- Claude Code v2.1.280+ lets you switch Opus 5.5 effort mid-session without breaking prompt cache — kimmonismus · 2026-09-23
- Everyone posts 3D render and SVG evals, but nobody shows the refusals and quotas — BLUECOW009 · 2026-09-23
- GPT-6 Astra vs Opus 5.5: Opus Wins Knowledge Work, OpenAI Keeps Hard STEM — johnseach · 2026-09-23
- Polymarket Odds: 91% Chance Claude 6 Ships by End of 2027 — Polymarket · 2026-09-23
- Five Hours of Heavy Opus 5.5 Use on /medium Burned Only 4% of the Weekly Limit — rudrank · 2026-09-23
- One Prompt, Full 3D Scene: GPT-6 Sol Needs Little Inference in 4-Model Test — rohanpaul_ai · 2026-09-23