Task is detecting null-finding claims, tested against multi-human-coded labels

RexDouglass · x · 2026-09-19

RexDouglass clarifies the task: whether an abstract claims a null finding. Their evaluation uses battle-tested instructions and multi-human-coded labels, giving a reliable ground truth.

Related event: Researchers Find Budget Open Models Struggle to Detect Null Findings, Newer Models Show Promise(9 posts)→

Original post →

More from Models

Models channel →