GDP.xlsx benchmark: 70 real spreadsheet tasks, best frontier agent scores only 38.3%
echen · x · 2026-10-01
A follow-up to GDP.pdf, the new GDP.xlsx benchmark simulates real-world spreadsheet work.
- 70 tasks across 12 domains; the hard part is decoding how a workbook works: which tab is stale, what red cells mean, which comment explains exceptions, which chart Finance trusts, checking termination clauses in Sheet 3.
- The best frontier agent scores 38.3% — reportedly Gemini 4 Argon.
The benchmark highlights implicit context and cross-sheet reasoning, which remain major weaknesses for current agents.
Related event: Surge AI Releases GDP.xlsx Benchmark; Top Agent Scores Just 38.3%(2 posts)→
More from Models
- Gemini 4 Scores Badly on cua-bench, Fueling Benchmaxxing Concerns — burny_tech · 2026-10-01
- Harvey's Legal Agent Benchmark shows conflicting scores: 19.6% vs 25.42% — 3scorciav · 2026-10-01
- Gemini 4 Argon reportedly live as Google's model cadence accelerates, unconfirmed — Dr_Singularity · 2026-10-01
- webAI's 3.6B TwIL-LM3-Pro beats VibeThinker-3B by 35% on formal logic, runs locally in 2GiB — rohanpaul_ai · 2026-10-01
- Quick benchmark: Sol 6.1 inference is nearly 6x slower than Opus despite better token efficiency — RexDouglass · 2026-10-01
- Reddit Users Question AA Intelligence Index After Sonnet 5.5 Outranks Fable 5.1 — Ill_Distribution8517 · 2026-10-01