Surge AI's DAYJOB benchmark: best agents pass only ~24% of real work
echen · x · 2026-09-24
- Surge AI launched DAYJOB, a benchmark suite for professional knowledge work, debuting with DAYJOB: Healthcare and DAYJOB: Finance — 130 expert-level assignments total.
- Motivation: GDPval-style tasks come with detailed instructions, but real work is a vague "@bob is that ready yet" — agents must reconstruct the task from spreadsheets, docs, and Slack threads and decide which sources to trust.
- Results: the strongest model passes only 24.7% on Healthcare and 23.9% on Finance — frontier agents remain far from surviving a 9 to 5.
More from Research
- One picture each: gradient descent and backprop explained without arithmetic — alfcnz · 2026-09-24
- Neural net series: multi-layer nets and why cross-entropy and squared error share the same gradient — alfcnz · 2026-09-24
- Teaching Thread: From Perceptrons to Hidden-Layer Representations in Driving — alfcnz · 2026-09-24
- Simulated fly on LSD: connectome wired with eyes shows T4/T5 activity up 10–20% — Merzmensch · 2026-09-24
- Anthropic's first biology lab result draws scrutiny: only ~4 parallel agents and no wet-lab details — basedjensen · 2026-09-24
- Predict rates, not yield: ML kinetic maps explain nickel cross-coupling failures — bravo_abad · 2026-09-24