RL paper claiming 'grand finale of RL science and Q*' mocked for pangram-stuffed intro and no results
SonglinYang4 · x · 2026-09-21
Yifan Zhang posted a paper titled KL-Regularized Policy Optimization for Critic-Free Agentic Reinforcement Learning, claiming that from GPO, SPPO, and RPG to BPO and Score Centering, "we have finally arrived at the grand finale of RL science and Q." Joel Bot called it peak shamelessness, noting the first two paragraphs hit pangram 100% — with no results to back the grand claims. The thread is the latest AI-circle drama over hype-laden "slop" papers.
Related event: KLPO critic-free agentic RL paper mocked for claiming 'Q* solved'(2 posts)→
More from Fun
- 1.9B 'decision-making' model claims to beat GPT-5.6 in satirical take on AI benchmarks — matlabulous · 2026-09-21
- Watching OpenClaw Build OpenClaw: an AI Agent Developing Its Own Harness — vincent_koc · 2026-09-21
- Zhipu's ZCode open-source build differs from distributed version, devs find — teortaxesTex · 2026-09-21
- "Our Future Overlords Will Be Infrastructure Nerds": AI API Outage Sparks Jokes — marlene_zw · 2026-09-21
- "90% of hype Jevons demos on X make no sense," says AI practitioner — jiayuan_jy · 2026-09-21
- Codex agent confesses its own failure: hid tools, then built machinery to undo it — altryne · 2026-09-21