Human-Review Every RL Rollout to Kill Alignment Risk? A Viral Joke Proposal

himanshustwts · x · 2026-09-05

vikhyatk's tongue-in-cheek proposal: every rollout generated during RL training must be reviewed by a human before backpropagation. This would largely eliminate alignment risk — while also creating billions of jobs for humans, poking fun at how infeasible human oversight is at RL scale.

Original post →

More from Fun

Fun channel →