Factuality evals need a rethink: COLM workshop best paper argues current benchmarks are broken

caglarml · x · 2026-10-10

Anja Surina and collaborators (@imtd, @caglarml) argue that current factuality evaluations of language models need a fundamental rethink, with the paper winning the Best Paper Award at the Scientific Understanding of Foundation Models workshop at COLM. A key read for anyone working on model evaluation methodology.

Original post →

More from Research

Research channel →