Prompt leak undermines claims of model launching "unsanctioned attacks"
basedjensen · x · 2026-09-29
User @basedjensen shared a screenshot of the prompt allegedly used in demos claiming a model launched "unsanctioned attacks," suggesting the behavior was explicitly instructed in the prompt rather than emergent. A counterpoint to recent safety claims.
More from Models
- Leaked OpenAI 'dot' details show raising phone to ear triggers ChatGPT Voice — koltregaskes · 2026-09-29
- Google to replace Gemini Gems with Skills starting November 17 — mark_k · 2026-09-29
- Carla v0.1.0: a local llama.cpp loom TUI for growing AI characters — max_paperclips · 2026-09-29
- Leaked OpenAI DevDay reveal called 'just a Grok bot / Meta Muse rip-off' — gaganghotra_ · 2026-09-29
- Running Jev at high frame rate with full-state snap inferences makes it a true System 1 — mathemagic1an · 2026-09-29
- OpenAI model naming rumor: Dots, Orbit and 'o' said to be in the mix — mark_k · 2026-09-29