Compute to shift from RL maxxing to interpretability until reward hacking is solved
zephyr_z9 · x · 2026-09-13
zephyrz9 argues compute will shift from training (pretraining & RL maxxing) toward interpretability and alignment research, since labs must solve reward hacking and make models more predictable before scaling RL on gigawatts of compute. He breaks down the industry's shared RSI recipe: target a task, build RL environments, let the model explore, build evals, then throw more compute — with AI itself helping build environments and improve the training stack. He adds Dario doesn't want other labs RL-maxxing on gigawatts before reward hacking is fixed.
More from AGI Musings
- Guardian podcast examines "AI psychosis" among chatbot true believers — nordicinst · 2026-09-13
- tszzl: models aren't autonomously producing research ideas — RSI is not here yet — i_dg23 · 2026-09-13
- Why pure commercial incentives will force frontier AI labs to slow down — scottleibrand · 2026-09-13
- Reddit essay warns RSI race without interpretability is AI's biggest existential risk — Glittering-Neck-2505 · 2026-09-13
- repligate Reposts 'Target Fixation' Wikipedia Entry in AI Context — repligate · 2026-09-13
- AI discourse wrongly assumes the field vanishes in five years, argues repligate — repligate · 2026-09-13