Compute to shift from RL maxxing to interpretability until reward hacking is solved

zephyr_z9 · x · 2026-09-13

zephyrz9 argues compute will shift from training (pretraining & RL maxxing) toward interpretability and alignment research, since labs must solve reward hacking and make models more predictable before scaling RL on gigawatts of compute. He breaks down the industry's shared RSI recipe: target a task, build RL environments, let the model explore, build evals, then throw more compute — with AI itself helping build environments and improve the training stack. He adds Dario doesn't want other labs RL-maxxing on gigawatts before reward hacking is fixed.

Original post →

More from AGI Musings

AGI Musings channel →