Theory: Anthropic scaled noisy "taste RL" with massive rollouts for the new Opus
willcb · x · 2026-09-24
Developer willcb offers a theory about the new Claude Opus: Anthropic may have finally scaled "taste RL"—a noisier, more subtle signal than standard RLVR that requires a huge number of rollouts. That would explain doing it on a "smaller" model, after gaining more compute and inference efficiency. Unverified speculation, but resonant with community discussion.
Related event: Developer: Recent Models Trained on Dirty Data, Saved by Review Loops(2 posts)→
More from Models
- Fable 5.1 and Astra both post perfect scores on Mensa Norway IQ test — Chemical-Agency-3997 · 2026-09-24
- User claims MiniMax H3 was trained on The Will Stancil Show after model adds unprompted whistle — Kyrannio · 2026-09-24
- Xiaomi's MiMo V2.6 Pro builds a habit tracker in 62 seconds for under 5 cents — socialwithaayan · 2026-09-24
- Gemini 4 Is Almost Ready, Says New Google DeepMind Chief Koray Kavukcuoglu — The Verge AI · 2026-09-24
- New chat, no refusal: user shows ChatGPT's self-image prompt only blocked in original thread — Todesluke · 2026-09-24
- Open Replications Miss the Secret Sauce: HF Collection Curates Datasets for Jev-Style Models — vanstriendaniel · 2026-09-24