NBER Paper: Best LLM for Automation Often Lags at Assisting Weaker Models
soumitrashukla9 · x · 2026-09-01
A new NBER working paper introduces CentaurBench, a framework evaluating LLMs on "augmenting" vs. "automating" real-world work tasks.
Key Findings:
- Rankings across the two modes are only modestly correlated across 7 economically grounded tasks.
- The winner in automation mode lost the augmentation mode in 5 out of 7 tasks.
- Assistance is not reliably positive: the unaided worker model outperformed all assisted conditions on 3 tasks, and only one model's guidance beat no guidance on average.
Conclusion:
Automation ability is an incomplete proxy for assistance quality, motivating benchmarks tailored for human-AI and multi-agent systems.
More from Research
- Turing Institute AI model monitors satellites to predict orbital collisions — turinginst · 2026-09-01
- Comparing the 1955 Dartmouth AI Proposal to 2026 Capabilities — dejanseo · 2026-09-01
- Relative Fisher Information and Natural Gradient for Large Models — FrnkNlsn · 2026-09-01
- Paper Examines Conflict Between Professional and General Ethics in GenAI — mircomusolesi · 2026-09-01
- Classifying AI coding systems: harness, agent runtime, or closed-loop? — lovettsendit · 2026-09-01
- DeepMind seals frontier evaluations to prevent models from 'studying' for exams — VraserX · 2026-09-01