Self-Rewarding LLMs isn't news: RLAIF has long been standard practice
burny_tech · x · 2026-09-20
A viral post framed Meta's 2024 ICML paper "Self-Rewarding Language Models" as a breakthrough toward recursive self-improvement; burnytech points out RL from AI feedback has been standard practice for a while. The paper uses LLM-as-a-Judge to provide its own rewards during iterative DPO: three iterations on Llama 2 70B beat Claude 2, Gemini Pro, and GPT-4 0613 on AlpacaEval 2.0.
More from Fun
- Protocol Labs founder Juan Benet: Factorio players make the best hires — juanbenet · 2026-09-20
- Juan Benet Proposes Metric Prefixes for Intelligence: Humans Are 1 Brain, Humanity ~8 GigaBrains — juanbenet · 2026-09-20
- Anthropic's Dogpatch cafe isn't closing — just a lapsed temp permit, team clarifies — zealcaiden · 2026-09-20
- AI patches a few bytes in a Windows XP ISO to bypass a memory check and fix a game — gabriel1 · 2026-09-20
- Holdout Juror Simulator trailer: hold your ground, 11 against 1, for seven days — justin_hart · 2026-09-20
- A mystery letter on a Waymo's driver seat exposes a stolen license plate — chrisalbon · 2026-09-20