DAYJOB Benchmark: Top Agents Pass Only ~24% of Real 9-to-5 Work Tasks
echen · x · 2026-09-29
Surge AI launched DAYJOB, a benchmark testing whether agents can survive a 9-to-5 job:
- Two environments — DAYJOB: Healthcare and DAYJOB: Finance — with 130 expert-built assignments modeled on real professional work
- Tasks are deliberately messy: short instructions, large environments, and agents must figure out what matters and carry work to completion
- The strongest models currently pass only 24.7% of Healthcare and 23.9% of Finance assignments
Takeaway: there's a large gap between completing a prompt and surviving a day at the office.
More from coding & agent
- Perplexity Agent API adds Profiles, Skills and managed connectors for reusable custom agents — AravSrinivas · 2026-09-29
- $25 credits: Droid ran 8 hours and finished, Devin burned out in under an hour — matanSF · 2026-09-29
- Stripe is building verification for AI agent web requests and user consent — jeff_weinstein · 2026-09-29
- Wonder Builds an Agent That Plans, Orders and Delivers All 21 Weekly Meals — LangChain · 2026-09-29
- React-to-Anything: stick likes to objects in video with Muse Spark + SAM 3.1 — nikhilaravi · 2026-09-29
- AI robotic integrator agent handles end-to-end MicroFactory deployment — ihorbeaver · 2026-09-29