Meta's Chief AI Officer Alexandr Wang: nobody knows how to solve alignment, favors AI watching AI
rohanpaul_ai · x · 2026-10-10
Meta Chief AI Officer Alexandr Wang told Cleo Abram that alignment remains one of the most open scientific questions in AI — nobody knows exactly how to solve it.
His proposed answer is scalable oversight: as models get smarter, a separate set of AI "watchers" observes them and keeps them in check. Since the watcher AIs must improve alongside the models they monitor, labs would need to keep building "smarter and smarter policing agents."
Wang noted Meta's Muse already implements a version of this, with a dedicated sentinel agent auditing what the main agent does.
More from AGI Musings
- Cosmos Institute founder warns AI 'pacing' regulators would gain near-unlimited power — luke_drago_ · 2026-10-10
- Musk says call center jobs will 'disappear fast' as the sector shrinks 4% a year — elonmusk · 2026-10-10
- Treat LLMs as 'weird little guys in your computer,' not software programs — dioscuri · 2026-10-10
- The irony: AI industry that promised to eliminate roles now can't hire them — pixlpa · 2026-10-10
- Patrick Collison: Personal AI Agents Will Reshape How Companies Exploit Consumer Bounded Rationality — scottleibrand · 2026-10-10
- iamtrask: the only moat is rare data, and math isn't rare data — iamtrask · 2026-10-10