Developer: Recent Models Trained on Dirty Data, Saved by Review Loops
Developer willcb argues recent models were trained on largely messy data but are now capable enough to fix issues within orchestrated workflows and review loops. He speculates the new Claude Opus scaled 'taste RL' beyond ordinary RLVR.
2026-09-24 ~ 2026-09-24 · 2 related posts
- Dev: Recent models were trained on broken data, saved by orchestration and review loops — willcb · 2026-09-24
- Theory: Anthropic scaled noisy "taste RL" with massive rollouts for the new Opus — willcb · 2026-09-24