We don't know how to train trustworthy AI models yet

JeffLadish · x · 2026-08-27

Jeff Ladish recommends an article by Grace Huckins arguing that no one knows how to train trustworthy AI models because we lack a real science of model motivations. While companies can harden environments, this won't solve deeper misalignment problems.

Original post →

More from Safety

Safety channel →