CentaurBench Reveals Strongest Models Aren't Always Best Assistants

soumitrashukla9 · x · 2026-08-27

UC Berkeley released CentaurBench, a new benchmark evaluating LLMs' "augmentation" capabilities—assisting a fixed worker model—versus traditional "automation" mode.

Key Findings:

Methodology:

Related event: CentaurBench: The Strongest Models Aren't Always the Best Assistants(3 posts)→

Original post →

More from Research

Research channel →