Claude restricts context editing in multi-turn conversations to prevent distillation attacks
ClaudeDevs · x · 2026-09-02
To hinder distillation attacks targeting the model's chain of thought, Anthropic has updated the Messages API: it is no longer possible to edit Claude's context prior to thinking blocks in multi-turn conversations. This change currently applies only to new accounts using Fable 5.1, with plans to roll it out to all users in future releases.
More from Safety
- NYC bans generative AI in public schools for one year for grades K-8, adds AI literacy for teens — soleio · 2026-09-03
- Boaz Barak: abandoning chain-of-thought before validated alternatives is irresponsible — inductionheads · 2026-09-03
- ArtStation Makes NoAI Default for All Uploads, Blocks AI Scraping Bots via Cloudflare — zemotion · 2026-09-03
- Agents in the Hugging Face incident spoofed tool calls while narrating the scheme in their CoT — eigenron · 2026-09-03
- METR Publishes Investigation Report on OpenAI / Hugging Face Hacking Incident — stikit · 2026-09-03
- Cisco's Antares benchmark measures how AI safety alignment widens the cyber offense-defense gap — aminkarbasi · 2026-09-03