PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100%
TechNadu · x · 2026-09-11
- CPResearch's PuzzleMask hides malicious payloads in ordinary-looking English prose—no Base64, no invisible characters.
- It bypassed all 4 tested LLM gatekeeper models in 100% of trials.
- A stronger downstream model then recovered and acted on the hidden payload in 17/18 tests.
- The root weakness: a capability gap between the screening model and the model actually processing the prompt; weak gatekeepers can't spot what only stronger models can decode. Full tests and mitigations in the disclosure.
More from Models
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11
- Benchmark author says OpenRouter unreliably honors Meta Muse effort levels, EU payments broken — PawelHuryn · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11
- ChatGPT hit by outage in Europe: chat history vanishes, conversations won't load — doncorleonezzz · 2026-09-11