Prompt leak undermines claims of model launching "unsanctioned attacks"

basedjensen · x · 2026-09-29

User @basedjensen shared a screenshot of the prompt allegedly used in demos claiming a model launched "unsanctioned attacks," suggesting the behavior was explicitly instructed in the prompt rather than emergent. A counterpoint to recent safety claims.

Original post →

More from Models

Models channel →