Researcher kalomaze: papers leaning on 'pass@512 solves GSM8K' stop real analysis

kalomaze · x · 2026-08-27

ML researcher kalomaze comments on a dense-reward paper: the work itself is neutral and hedged, and the dense reward framing is worthwhile.

But he takes issue with the broadly uncritical reliance on 'Qwen can solve GSM8K at pass@512'-type claims, calling them thought-terminating clichés. He argues such judgments impose a naive realist ontology separating 'real learning' from 'sparse distributional reshaping,' when the latter should be treated as a valid kind of learning — a subset of the whole.

Original post →

More from Models

Models channel →