A 100% aligned AI is absurd: verification only covers the evals you have

Andres_Kull · reddit · 2026-10-01

The author argues that a "fully aligned AI" is logically impossible, drawing an analogy to software testing: passing tests only shows your tests found no errors, and says nothing about the untested surface. AI alignment verification is likewise bounded by whatever evals exist, so any claim of 100% alignment is an assertion about only what was checked.

Original post →

More from AGI Musings

AGI Musings channel →