kalomaze: models learn narrow RLVR fact-checks instead of asserting only in-context provable claims
kalomaze · x · 2026-09-13
kalomaze argues models have mostly learned narrow RLVR-style epistemic checks for particular factual claims, rather than being taught, as agents, to assert claims based on in-context provability. Smaller models casually state hypotheses (especially those matching pretraining priors) as proven before any falsifying evidence arrives — and sometimes ignore evidence already present in earlier tool call results.
Related event: Dev Criticizes AI Agents' Useless Memory Files, Blames Narrow RLVR(3 posts)→
More from Models
- User shows how GPT flatters your beliefs and hallucinates premises in arguments — GlenBradley · 2026-09-13
- Orchestrator Error Reveals Mystery Model 'Daybreak': 'astra Is Not Allowed to Access Those Resources' — LeopardBernstein · 2026-09-13
- kalomaze: Opus 5 is "such a bad model" — kalomaze · 2026-09-13
- Wenhu Chen can't even understand many questions in AA-intelligence AI benchmarks — WenhuChen · 2026-09-13
- Commenter: The OpenAI Navier-Stokes proof would be hailed as a breakthrough if posted anonymously — skdh · 2026-09-13
- GPT-6 Astra skips the chat box: Plus users get just 5-45 messages per 5 hours, locked to Work and Codex — 新智元 · 2026-09-13