Self-Rewarding LLMs isn't news: RLAIF has long been standard practice

burny_tech · x · 2026-09-20

A viral post framed Meta's 2024 ICML paper "Self-Rewarding Language Models" as a breakthrough toward recursive self-improvement; burnytech points out RL from AI feedback has been standard practice for a while. The paper uses LLM-as-a-Judge to provide its own rewards during iterative DPO: three iterations on Llama 2 70B beat Claude 2, Gemini Pro, and GPT-4 0613 on AlpacaEval 2.0.

Original post →

More from Fun

Fun channel →