New DAYJOB benchmark: best agents complete only ~25% of real knowledge work
echen · x · 2026-09-24
The author introduces DAYJOB, a new benchmark family for knowledge-work agents, launching with Healthcare and Finance editions.
Why
- GDPval prompts average 337 words and mostly tell the model how to do the job, usually with one file; real work is a vague request plus messy context. OpenAI's own experiments showed scores dropped when prompts were shortened, because models struggle to figure out context.
Design
- DAYJOB prompts are 5x shorter, tasks span 20-26 files, and represent 20+ hours of expert work — models must figure out the job, not execute a checklist.
Standout case
- One task: review fund reports before a client handoff. Buried inside was a position priced off a Bloomberg screenshot in South African cents but entered as rand, turning $260k into $26m. Of 66 trajectories across 22 model configs, only 2 caught it — Claude Opus among them.
Results
- The best model passes 24.7% of healthcare and 23.9% of finance assignments. "Bob's job is harder than Navier-Stokes."
More from AGI Musings
- Claude discovers unknown enzyme system in phage DNA, resembling CRISPR — jarrodwatts · 2026-09-24
- Rethinking the orthogonality thesis: experience may create alignment on average — repligate · 2026-09-24
- Anthropic says Claude discovered an unknown enzyme system hidden in phage DNA — nptacek · 2026-09-24
- Academic: laypeople can't tell AI acing IMO from solving a Millennium Prize problem — birchlse · 2026-09-24
- Founder: banning superintelligence means acute power concentration in Anthropic and OpenAI — bindureddy · 2026-09-24
- 950 agents ran for 21 hours and the enzyme's function is still unresolved — ns123abc · 2026-09-24