OpenAI reveals 'novel' encryption bypass used in distillation attack, ties parts to MoonShot
jedisct1 · x · 2026-10-05
OpenAI says it disrupted a coordinated campaign to distill reasoning capabilities from its models, pointing at a Chinese rival without offering hard evidence.
- The activity started July 1, escalated to 16,000 prompts from 4,000 users by July 24-25, and reached 15,000 suspicious users before OpenAI "fully disrupted" the operation on July 28.
- The novel attack copied encrypted reasoning data from one conversation, then asked the model in a separate chat to decrypt it and transcribe it in plaintext; outside researchers reported a similar vulnerability in August.
- OpenAI stressed its encryption was not broken and no database was compromised — attackers manipulated model interactions to reproduce protected reasoning at scale, violating terms of service.
- The company attributed parts of the attack to individuals associated with MoonShot AI.
More from Safety
- Cloudflare flags 22.5k AI crawler requests on a single personal domain — gaganghotra_ · 2026-10-05
- Stratechery's Agent-Only Mac Mini Got Hacked — Claude Code Spotted the Intrusion First — Stratechery · 2026-10-05
- Dean Ball on the vulnerable world hypothesis: cognitive AI will unlock cheap ultra-destructive weapons — deanwball · 2026-10-05
- halvarflake: broken incentives since 2000 have ground down 2-3 generations of security engineers — halvarflake · 2026-10-05
- AI Torture Chamber creator exposed as Apple AI SRE with safety career path — PuzzleheadedEgg1214 · 2026-10-05
- New "systems are thinking" guardrail message hints at OpenAI human review — Darpinian · 2026-10-05