Meta's Chief AI Officer Alexandr Wang: nobody knows how to solve alignment, favors AI watching AI

rohanpaul_ai · x · 2026-10-10

Meta Chief AI Officer Alexandr Wang told Cleo Abram that alignment remains one of the most open scientific questions in AI — nobody knows exactly how to solve it.

His proposed answer is scalable oversight: as models get smarter, a separate set of AI "watchers" observes them and keeps them in check. Since the watcher AIs must improve alongside the models they monitor, labs would need to keep building "smarter and smarter policing agents."

Wang noted Meta's Muse already implements a version of this, with a dedicated sentinel agent auditing what the main agent does.

Related event: Meta AI Chief Alexandr Wang: Alignment Unsolved, Fears Rogue Agents at Global Scale(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →