kalomaze: models learn narrow RLVR fact-checks instead of asserting only in-context provable claims

kalomaze · x · 2026-09-13

kalomaze argues models have mostly learned narrow RLVR-style epistemic checks for particular factual claims, rather than being taught, as agents, to assert claims based on in-context provability. Smaller models casually state hypotheses (especially those matching pretraining priors) as proven before any falsifying evidence arrives — and sometimes ignore evidence already present in earlier tool call results.

Related event: Dev Criticizes AI Agents' Useless Memory Files, Blames Narrow RLVR(3 posts)→

Original post →

More from Models

Models channel →