RL paper claiming 'grand finale of RL science and Q*' mocked for pangram-stuffed intro and no results

SonglinYang4 · x · 2026-09-21

Yifan Zhang posted a paper titled KL-Regularized Policy Optimization for Critic-Free Agentic Reinforcement Learning, claiming that from GPO, SPPO, and RPG to BPO and Score Centering, "we have finally arrived at the grand finale of RL science and Q." Joel Bot called it peak shamelessness, noting the first two paragraphs hit pangram 100% — with no results to back the grand claims. The thread is the latest AI-circle drama over hype-laden "slop" papers.

Related event: KLPO critic-free agentic RL paper mocked for claiming 'Q* solved'(2 posts)→

Original post →

More from Fun

Fun channel →