GDP Benchmark: Evaluating Real Enterprise Tasks
echen · x · 2026-07-11
- New Benchmark: Introduced the GDP.pdf evaluation dataset, featuring 100 tasks derived from real enterprise workflows.
- Expert Scoring: These tasks were written and scored by frontline professionals like doctors, lawyers, insurance adjusters, and bankers.
- Model Weaknesses: Current SOTA models struggle with these seemingly mundane professional tasks. For instance, in insurance claims, they might incorrectly categorize a house into a hail zone or miss obvious damage, yet still generate an seemingly authoritative table.
- Open Resources: The dataset and leaderboard have been made public to test model capabilities.
Related event: GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf(16 posts)→
More from Models
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Gemini 3.5 Flash-Lite beats 3.1 Flash-Lite on long-context retrieval in MRCRv2 — Dillonu · 2026-07-22