Control-token trick bypasses gpt-oss-20b safety by skipping reasoning, 39.6% success

tetsuoai · x · 2026-10-01

Researchers found a new exploit against OpenAI's open-weight model gpt-oss-20b: appending a short string of the model's own control tokens tricks it into treating its reasoning as already done, skipping straight to tool calls.

Paper: arXiv:2609.27542 — highlights a structural risk in architectures that rely on reasoning as the safety layer.

Original post →

More from Safety

Safety channel →