Factuality evals need a rethink: COLM workshop best paper argues current benchmarks are broken
caglarml · x · 2026-10-10
Anja Surina and collaborators (@imtd, @caglarml) argue that current factuality evaluations of language models need a fundamental rethink, with the paper winning the Best Paper Award at the Scientific Understanding of Foundation Models workshop at COLM. A key read for anyone working on model evaluation methodology.
More from Research
- 20-minute explainer breaks down how Tesla trains FSD: 8 cameras, 36fps, 2B signals to 2 outputs — PTrubey · 2026-10-10
- 21 researchers release white paper on Visual General Intelligence as a path to AGI — HirokatuKataoka · 2026-10-10
- Epoch AI: Over half of arXiv math papers in 3 subfields now acknowledge AI use — burny_tech · 2026-10-10
- Researcher faults standard Transformer diagrams for hiding self-attention as a black box — PlisSergey · 2026-10-10
- Xiaomi's MiMo-V2.6: agents run their own RL loop, DeepSWE score hits 72.6 for $2.6M — rohanpaul_ai · 2026-10-10
- New fastest deterministic 3SUM algorithm hits n^1.9961, matching randomized bound — basedjensen · 2026-10-10