Study finds most accurate models are not always the most preferred by users
soumitrashukla9 · x · 2026-08-27
A highlight of work led by Mina Lee finds that the most accurate models are not always the ones human users prefer the most.
For instance, Opus and Sonnet are great at performing tasks on their own but not in providing assistance, while GPT-5-Mini was strong in both dimensions. Gemini models were found to be stronger as assistants than as automators.
Related event: CentaurBench: The Strongest Models Aren't Always the Best Assistants(3 posts)→
More from Research
- Scripps Research wins $19.5M NSF grant for AI-powered autonomous chemistry lab — CatAstro_Piyush · 2026-08-27
- Anatomy of Company Brains: 4 shared components across 9 projects — femke_plantinga · 2026-08-27
- 411k cut-outs from Britannica: 29M param model segments historical illustrations — vanstriendaniel · 2026-08-27
- van der Schaar Lab: What gets hidden when medicine is built around the average patient? — MihaelaVDS · 2026-08-27
- VGGT-SLAM++: Complete Visual SLAM System with Sim(3) Backend — rsasaki0109 · 2026-08-27
- Researcher kalomaze: papers leaning on 'pass@512 solves GSM8K' stop real analysis — kalomaze · 2026-08-27