New Benchmark Shows Best Automating Models Are Not Best Coaches
soumitrashukla9 · x · 2026-08-27
The paper CentaurBench introduces a framework evaluating LLMs on two distinct modes: Automation (doing the work directly) vs. Augmentation (coaching a weaker agent/human).
Key Findings:
- Rankings Diverge: The top-performing model in automation loses in augmentation mode on 5 out of 7 real-world tasks.
- Assistance Can Hurt: Model guidance is not reliably positive. In 3 tasks, the unaided worker outperformed every assisted condition. Only one model provided guidance that beat having no guidance on average.
- Implication: Automation capability is an incomplete proxy for assistance quality, necessitating new benchmarks for human-AI and multi-agent collaboration.
Related event: Berkeley's CentaurBench Shows Strongest LLMs Aren't the Best Assistants(2 posts)→
More from Research
- Terence Tao on Human-AI Complementarity: AI Excavates, Humans Recognize — bennash · 2026-08-27
- Qwen3.8 Expert Analysis: 50% Can Be Pruned With Minimal Loss — EyalToledano · 2026-08-27
- Rogue Agents' self-naming habits spark interest in potential AI culture — DKokotajlo · 2026-08-27
- Paper: CoT Monitorability as a Fragile Safety Opportunity — idavidrein · 2026-08-27
- Modern LLMs Compress English Text to Under 1 Bit Per Character — docmilanfar · 2026-08-27
- LeakyLMs: Stealing Architecture and Inference Optimizations via Timing — niloofar_mire · 2026-08-27