Mechanism design researcher Andrew Koh joins DeepMind with new AI alignment framework

burny_tech · x · 2026-09-04

Andrew Koh announced he's joining Google DeepMind, releasing a mechanism design framework for AI alignment and control. Largely conceptual, it offers stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a unified way to value alignment, interpretability, capability, and control.

Original post →

More from Companies & People

Companies & People channel →