New Benchmark Shows Best Automating Models Are Not Best Coaches

soumitrashukla9 · x · 2026-08-27

The paper CentaurBench introduces a framework evaluating LLMs on two distinct modes: Automation (doing the work directly) vs. Augmentation (coaching a weaker agent/human).

Key Findings:

Related event: Berkeley's CentaurBench Shows Strongest LLMs Aren't the Best Assistants(2 posts)→

Original post →

More from Research

Research channel →