Researchers Find Budget Open Models Struggle to Detect Null Findings, Newer Models Show Promise

Researchers jonmellon and RexDouglass shared hands-on lessons on X from using LLMs to classify academic abstracts: the task was to determine whether a paper's abstract reports a null finding, with the benchmark built on well-polished annotation instructions and labels hand-coded by multiple annotators. The results: most cheap models performed mediocrely on this task, while one new model showed promise for the first time.

Confirmed

Why it matters

2026-09-19 ~ 2026-09-19 · 9 related posts

Primary sources