CentaurBench Reveals Strongest Models Aren't Always Best Assistants
soumitrashukla9 · x · 2026-08-27
UC Berkeley released CentaurBench, a new benchmark evaluating LLMs' "augmentation" capabilities—assisting a fixed worker model—versus traditional "automation" mode.
Key Findings:
- Model rankings differ significantly between automation (direct execution) and augmentation (providing assistance).
- For instance, Opus and Sonnet excel at automation but lag in assistance; Gemini models are stronger assistants than automators; GPT-5-Mini performs well in both.
Methodology:
- Evaluated across 7 real-world economic tasks derived from GDPval and Anthropic's Economic Index.
- Outputs are scored blindly using a validated LLM-as-judge framework.
- Top augmenters share three traits: defining what constitutes "good" work, structuring reasoning steps, and reinforcing task requirements.
Related event: CentaurBench: The Strongest Models Aren't Always the Best Assistants(3 posts)→
More from Research
- Miles now supports RL training for Qwen and GLM models with high-performance kernels — ying11231 · 2026-08-28
- Miles-diffusion introduces LoRA SFT for fast post-training of diffusion models — ying11231 · 2026-08-28
- SovietRxiv adds 7,000 translated Soviet scientific papers to archive — generativist · 2026-08-28
- 87% of "Quantum Supremacy" Claims Fail Under Real-World Testing, Physicist Says — AryHHAry · 2026-08-28
- FP-AMB: a first-person agent memory benchmark that tells you why each miss happened — LowDistribution3995 · 2026-08-28
- Penn & Yale Paper: Conformal Prediction Calibrations Diverge — ReCal Makes Them Reproducible — burkov · 2026-08-28