Burkov on RL Whack-a-Mole: Fix One Failure, Another Pops Up
burkov · x · 2026-09-12
Drawing on his own experience maintaining classifiers, Andriy Burkov argues RLHF is whack-a-mole: patch one issue and another appears. He predicts reinforcement learning will elicit longer responses but cause failures elsewhere.
More from Models
- Arcee open-sources NAC agent harness; 489 tool calls, 67 passing tests on DeepSeek V4.1 Flash — MaziyarPanahi · 2026-09-12
- Ant's Ling-3.0-flash-VL Scores 25 on AA Index With Just 5.5B Active Params — ArtificialAnlys · 2026-09-12
- Open-weight competition from China blocks AI labs from Uber-style price hikes, argues developer — firasd · 2026-09-12
- Frontier labs' privacy terms are 'insane': one thumbs-up can void your chat protections — niloofar_mire · 2026-09-12
- Redditor argues Astra hides chain of thought to prevent distillation, not to improve the model — bfkill · 2026-09-12
- Cursor ships CursorBench 4.0; speculation swirls that Gemini 3.8 post-training differs — ivan_bezdomny · 2026-09-12