Anthropic red team warns a downloadable Chinese AI model can now build working hacks on its own
ross2000 · reddit · 2026-09-30
Anthropic's Frontier Red Team warns that a Chinese AI model anyone can download can now autonomously build working hacks. The Reddit post flags it as worrying news, though it links no further details — the core claim is an open-weights model crossing a cyber-offense capability threshold per Anthropic's own red team assessment.
More from Safety
- Ask HN-style: How to Securely Scale ChatGPT Desktop + Blender MCP Company-Wide — 009fe3 · 2026-09-30
- America.gov chatbot refuses queries containing anything vaguely resembling PII — zsakib_ · 2026-09-30
- AI CEOs sign White House Accord on Superintelligence, Trump calls it 'morally binding' — ShakeelHashim · 2026-09-30
- Washington Post investigation exposes misuse of Flock Safety's nationwide license plate reader network — ScottNover · 2026-09-30
- OpenAI Faces Lawsuit Over the Hugging Face Hack — wiredmagazine · 2026-09-30
- Trump doubles down on light-touch AI regulation, says industry self-regulation suffices — pstAsiatech · 2026-09-30