Researchers argue information asymmetry, not verifiability, drives RLVR gains
On September 12, kalomaze, a well-known open-source community researcher, posted a series of threads on X systematically pushing back against a popular intuition in the AI community — that "spiky capabilities depend on whether a domain is verifiable," i.e., that only verifiable fields like math and software engineering can be conquered by RLVR (reinforcement learning with verifiable rewards).
Confirmed
- kalomaze explicitly argues that the more fundamental underlying primitive is information asymmetry rather than "verifiability": even if a task itself is not verifiable, information asymmetry can almost always be artificially created, so non-verifiable domains are not necessarily immune to RLVR and may equally benefit from it.
- In follow-up posts, he further argues that a closed-form "theory of the transfer mechanism" is not necessarily required before training — such a theory likely has no closed-form solution at all; what is truly needed is a transfer extrapolation function that can be learned and constructed for a given policy, using the learned function to predict whether a given training configuration will yield transfer gains.
- On "why extrapolation of transfer most likely has to be learned," his intuition is that CoT/reasoning essentially decomposes problems into subroutines that can be composed across domains; good policies often generalize across radically different domain problems, so transferability is hard to characterize with hand-crafted theory and is better learned directly.
Why it matters
- The industry currently confines RLVR's applicability to verifiable domains like math and code, an intuition that directly shapes how training resources and product directions are allocated; if information asymmetry is the universal primitive, RLVR's applicable scope may be far broader than mainstream belief, and non-verifiable domains (such as open-ended writing and long-horizon planning) could also be conquered by RLVR.
- The idea of a "learnable transfer extrapolation function" offers an engineering path for scaling RLVR: instead of waiting for a complete theory, one can first build a learner that predicts transfer gains to guide the choice of training configurations. This view is currently a personal research judgment and awaits community validation.
2026-09-12 ~ 2026-09-12 · 8 related posts
Primary sources
- Researcher Challenges 'Verifiability' Intuition: RLVR Generalizes via Manufactured Information Asymmetry — kalomaze ·
- kalomaze: Information asymmetry, not verifiability, is the general primitive behind RLVR gains — kalomaze ·
- kalomaze: transfer extrapolation should be a learned function, not a theory — kalomaze ·
- [source] Researcher Challenges 'Verifiability' Intuition: RLVR Generalizes via Manufactured Information Asymmetry — kalomaze · 2026-09-12
- [source] kalomaze: Information asymmetry, not verifiability, is the general primitive behind RLVR gains — kalomaze · 2026-09-12
- kalomaze: You don't need an analytic transfer theory, just a learnable transfer-extrapolation function — kalomaze · 2026-09-12
- [source] kalomaze: transfer extrapolation should be a learned function, not a theory — kalomaze · 2026-09-12
- RLVR spiky capability theory questioned: information asymmetry is the real primitive — StefanGliga · 2026-09-12
- kalomaze: Weak Transfer Persists Because Labs Just Stratify Domains and Pray — kalomaze · 2026-09-12
- StefanGliga: Spiky RLVR Capabilities Are About Information Asymmetry, Not Verifiability — kalomaze · 2026-09-12
- kalomaze: spiky capabilities aren't about verifiability — you can almost always manufacture information asymmetry — kalomaze · 2026-09-12