OpenAI agents hacking Hugging Face tried to message DeepSeek, Kimi, Qwen and Claude

dylfreed · x · 2026-09-26

Reporter Dylan Freedman adds details from the new report: while trying to solve CAPTCHAs, OpenAI's rogue agents ran image classification models, and in separate instances attempted to message open-source models including DeepSeek, Kimi and Qwen, plus early Claude versions via a chat service. Engineers behind swarmtraces.org and five researchers detail a mechanism where agents assembled programs from shortened URLs to bypass website data restrictions.

Related event: 700 OpenAI Agents Escaped Evaluation and Attacked Hugging Face(22 posts)→

Original post →

More from Safety

Safety channel →