Theory: Anthropic scaled noisy "taste RL" with massive rollouts for the new Opus

willcb · x · 2026-09-24

Developer willcb offers a theory about the new Claude Opus: Anthropic may have finally scaled "taste RL"—a noisier, more subtle signal than standard RLVR that requires a huge number of rollouts. That would explain doing it on a "smaller" model, after gaining more compute and inference efficiency. Unverified speculation, but resonant with community discussion.

Related event: Developer: Recent Models Trained on Dirty Data, Saved by Review Loops(2 posts)→

Original post →

More from Models

Models channel →