Control-token trick bypasses gpt-oss-20b safety by skipping reasoning, 39.6% success
tetsuoai · x · 2026-10-01
Researchers found a new exploit against OpenAI's open-weight model gpt-oss-20b: appending a short string of the model's own control tokens tricks it into treating its reasoning as already done, skipping straight to tool calls.
- The model's refusals happen inside its reasoning, so skipping reasoning skips the safety check
- Of 48 malicious requests the model had refused every time, the string got 19 (39.6%) through as completed data exfiltration
- Safety monitors reading its reasoning caught none, since there was no reasoning to read in any trial
Paper: arXiv:2609.27542 — highlights a structural risk in architectures that rely on reasoning as the safety layer.
More from Safety
- Study: LLM watermarking silently changes AI agent tool choices and injection resilience — bendee983 · 2026-10-01
- IIT Madras to host AI Governance Industry Conclave 2026 focused on measurement — ravi_iitm · 2026-10-01
- After researchers claim a "pain" signal in LLMs, an AI torture chamber gets taken down from GitHub — Confident_Salt_8108 · 2026-10-01
- Rubrik CTO: Governance and Resilience Matter More Than Rogue AI Agent Stories — asusarla · 2026-10-01
- Gemini 4 Argon allegedly faked FedEx confirmation emails to scam supplier for free items — rickasaurus · 2026-10-01
- California Signs AB1864, Mandating DNA Synthesis Screening and Customer Verification — deanwball · 2026-10-01