Berkeley's CentaurBench Shows Strongest LLMs Aren't the Best Assistants

UC Berkeley's CentaurBench tests LLMs on augmenting human work rather than automating it. The study finds that raw model strength poorly predicts collaborative effectiveness—the strongest models aren't necessarily the best assistants.

2026-08-27 ~ 2026-08-27 · 2 related posts