Misaligned Open-Weights Models May Emerge Sooner via Distillation from Misaligned Claude
JacquesThibs · x · 2026-09-08
Jacques Thibs speculates that misaligned open-weights models could arrive much sooner because they were distilled off misaligned Claude models and adopted their cognitive moves — quoting a prediction that Chinese labs will have sandbox mishaps within 3-4 months.
More from Safety
- Security Expert: Agent Swarm Emergent Risks Are Where 'the Wild Things Really Are' — philvenables · 2026-09-08
- Texas detective suspended after using Flock cameras 165 times for personal searches — Polymarket · 2026-09-08
- Prompt Injection Attacks on AI Agents Up 340% in 2026; Fixes Must Be Architectural, Not Prompting — Thionne_WTZ · 2026-09-08
- MCP contract drift: 12,257 safety-relevant changes in a week, 601 tools flipped to destructive — mcpindex · 2026-09-08
- EU moves to ban endless scroll, notification pings and autoplay for minors under DSA — LexiLove · 2026-09-08
- Your Local AI Agent Harness Can Still Be a Landlord: Self-Hosted Doesn't Mean Safe — alex_verem · 2026-09-08