AI safety assessments should move upstream to training runs
DavidSKrueger · x · 2026-07-22
- A quoted post argues that third-party assessments need to move earlier in the lifecycle, especially to training-run assessments.
- The key claim is that external deployment is no longer the main harm barrier: a sufficiently capable misaligned model could already cause damage during training or internal evaluations.
Related event: METR: AI Models Pose Risks Even Before Public Deployment(2 posts)→
More from AGI Musings
- Agentic breakouts split into stochastic failures and adversarial abuse — danielrock · 2026-07-22
- Age of Subjectivity argues complexity depends on the observer, not just the system — drmichaellevin · 2026-07-22
- Podcast discusses emulated minds that could share lifetimes in seconds — juanbenet · 2026-07-22
- AI slop detectors are useless, the post argues, because they rely on AI and the same training data — iamKierraD · 2026-07-22
- A post on the loss of hacker culture says the real issue is an instinct to obey — fkasummer · 2026-07-22
- Critics warn iterative deployment raises the stakes after every frontier AI failure — DavidSKrueger · 2026-07-22