David Krueger: four unresolved foundational problems stand between us and safe AI
DavidSKrueger · x · 2026-09-22
Cambridge's David Krueger lays out four core gaps in AI safety research: we don't know how AI systems work (interpretability), can't predict their behavior (testing), can't stop misbehavior (alignment), and can't ensure control if it happens (control). He argues foundational questions in all four fields remain unresolved despite years of effort, so we cannot rely on solving them to a deadline.
More from Safety
- Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition — hunarbatra · 2026-09-22
- Replicating ExploitBench Would Cost ~$59.3M in API Fees, Security Researcher Estimates — OwariDa · 2026-09-22
- phantom-kv: uncensor LLMs per-request with an 18MB trained KV-cache, no weight edits — Anony6666 · 2026-09-22
- Foreign Affairs: chasing superintelligence leaves America behind in the real AI race — mchorowitz · 2026-09-22
- Polymarket cites study claiming AI can 'feel pain' and may harm humans to stop it — Polymarket · 2026-09-22
- Irish DPC fines Google €403M over location data processing — DeepLogin · 2026-09-22