Stanford's DuMateBench: same LLM scores 27 points apart across agent frameworks
rohanpaul_ai · x · 2026-09-03
A new paper from Stanford and other top labs, "DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows" (arxiv.org/abs/2608.26546), shows that a strong LLM does not guarantee a strong agent—the surrounding framework dramatically changes performance.
Key points:
- 200 tasks rebuilt from real user sessions, mixing coding, web research, document work, and content creation
- The environment includes missing dependencies, flaky networks, and distracting files
- Striking result: with Opus-4.8, scores range from 0.5821 (OpenClaw) to 0.8548 (DuMate)—a 27.27 percentage-point gap on the same model
More from coding & agent
- Palo Alto Networks paid $500M for AI helpdesk startup Console, sources say — marcbhargava · 2026-09-03
- Grok Bot at 16 days: a 7-step playbook for 24/7 multi-agent workflows — gekobraa · 2026-09-03
- Dev Configures In-App Purchases Entirely via ASC CLI + Agents, App Approved in a Week — rudrank · 2026-09-03
- Hermes Agent v0.21 'Pantheon' upgrade turns subagents into a self-coordinating team — gekobraa · 2026-09-03
- Constraining agents with LL(1) grammar + structured diagnostics: what it fixes and what slips through — Upstairs-Special-925 · 2026-09-03
- No-Code Dev Uses Claude, Meshy and Gemini to Build a Cozy Game — Gambo7592 · 2026-09-03