Will Brown: robustly scaling reward modeling is the key problem for capabilities and safety

willcb · x · 2026-09-17

Will Brown reaffirms his earlier point that the most important problem for both capabilities and safety is robustly scaling reward modeling to arbitrary soft attributes that are hard to deterministically verify.

Original post →

More from Research

Research channel →