Burkov on RL Whack-a-Mole: Fix One Failure, Another Pops Up

burkov · x · 2026-09-12

Drawing on his own experience maintaining classifiers, Andriy Burkov argues RLHF is whack-a-mole: patch one issue and another appears. He predicts reinforcement learning will elicit longer responses but cause failures elsewhere.

Original post →

More from Models

Models channel →