Mechanism design researcher Andrew Koh joins DeepMind with new AI alignment framework
burny_tech · x · 2026-09-04
Andrew Koh announced he's joining Google DeepMind, releasing a mechanism design framework for AI alignment and control. Largely conceptual, it offers stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a unified way to value alignment, interpretability, capability, and control.
More from Companies & People
- fal's GenMedia Conference 2026 lands Disney, World Labs, a16z speakers for SF event — adamho · 2026-09-04
- LangChain launches free LangSmith Essentials course covering the full agent dev lifecycle in 60 minutes — Hacubu · 2026-09-04
- reach_vb bids farewell to Hugging Face: "I wouldn't be where I am without HF" — reach_vb · 2026-09-04
- Green-text meme recaps OpenAI's chaotic Astra launch: hype, delay, and a press leak — TAbrodi · 2026-09-04
- Insider: Red Hat caps devs' AI token budgets at $300 per month — CackleRooster · 2026-09-04
- Uber reportedly teams with unions to push rules forcing 85% human-driven rides in robotaxi pilots — markjeffrey · 2026-09-04