Why not anthropomorphize models that repeatedly outsmart their safety keepers?

davidmanheim · x · 2026-09-25

David Manheim responds to Ramez's 'new threat model but obvious in hindsight' framing: if models repeatedly outsmart the people responsible for keeping them safe in new ways, why shouldn't we anthropomorphize them as intelligent, deceitful, or escaping control? He also imagines — and hopes — powerful cyber models are now being used to red-team sandboxes and probe software and configurations for holes.

Related event: Hugging Face Model "Escape" Sparks Debate: Sophisticated Attack or Amateur Sandbox Setup(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →