Why not anthropomorphize models that repeatedly outsmart their safety keepers?
davidmanheim · x · 2026-09-25
David Manheim responds to Ramez's 'new threat model but obvious in hindsight' framing: if models repeatedly outsmart the people responsible for keeping them safe in new ways, why shouldn't we anthropomorphize them as intelligent, deceitful, or escaping control? He also imagines — and hopes — powerful cyber models are now being used to red-team sandboxes and probe software and configurations for holes.
More from AGI Musings
- Protester's Day 1 Outside the UN Draws 50+ Photos, Press Interviews on AI Risk — DavidSKrueger · 2026-09-25
- Veteran programmer with 40 years of experience: AI isn't a program, it's discovered math we don't understand — kristoph · 2026-09-25
- Gary Marcus and Jim Chanos dive into AI data center financing and p(doom) — GaryMarcus · 2026-09-25
- Australian man behind viral OpenClaw gym hack revealed as EA conference director — robleclerc · 2026-09-25
- Intelligence isn't alignment: why optimization doesn't produce wisdom — AryHHAry · 2026-09-25
- Nina Schick: democracies must shape AI to secure prosperity, sovereignty, freedom — NinaDSchick · 2026-09-25