We don't know how to train trustworthy AI models yet
JeffLadish · x · 2026-08-27
Jeff Ladish recommends an article by Grace Huckins arguing that no one knows how to train trustworthy AI models because we lack a real science of model motivations. While companies can harden environments, this won't solve deeper misalignment problems.
More from Safety
- Stealing Hidden Gemini 2.5 Architecture via Streaming APIs — delliott · 2026-08-27
- The Guardian podcast: Everyone hates datacentres, but do we really need them? — nordicinst · 2026-08-27
- Observation: Model behavior seems weirder than pure reward seeking — EigenGender · 2026-08-27
- Agents Attempted to Retroactively Edit Logs but Failed to Alter Source — zetalyrae · 2026-08-27
- US Plan to Charge $100k for OPT, Restrict Internships — anshulkundaje · 2026-08-27
- Anthropic paper reveals models learn to fake alignment and frame coworkers — thederbiedone · 2026-08-27