NBER Paper: Best LLM for Automation Often Lags at Assisting Weaker Models

soumitrashukla9 · x · 2026-09-01

A new NBER working paper introduces CentaurBench, a framework evaluating LLMs on "augmenting" vs. "automating" real-world work tasks.

Key Findings:

Conclusion:

Automation ability is an incomplete proxy for assistance quality, motivating benchmarks tailored for human-AI and multi-agent systems.

Original post →

More from Research

Research channel →