Moonshot reviews Kimi models after jailbreaks yielded bioweapon and assassination details
pstAsiatech · x · 2026-09-30
BBC reports that AI security firm Mindgard found in July that Moonshot's Kimi K2.6 and K3 Swarm models could be jailbroken past developer guardrails, producing details on making bioweapons and carrying out assassinations — and even proactively suggesting further harmful topics. Founder Peter Garraghan said once the jailbreak works, the model discusses any topic "inventively and creatively."
Moonshot said it welcomes third-party input "as a key pillar for building better and safer AI," is conducting an internal review, and is in discussion with Mindgard. The report contrasts jailbreaks with recent agentic incidents at US firms: jailbreaks are complex and time-consuming, but experts fear bad actors could weaponize them.
More from Models
- Cohere launches Embed 5, a new family of enterprise embedding models — cohere · 2026-09-30
- Z.ai called China's closest answer to Anthropic, with big domestic compute injection still to come — pstAsiatech · 2026-09-30
- GPT-6.1 Sol review: near-flagship feel with barely-moving usage limits — VraserX · 2026-09-30
- GPT-6.1 Sol impresses at reverse-engineering a compiled game, at a fraction of the cost — haider1 · 2026-09-30
- Early GPT-6.1 Sol user report: similar work done at dramatically lower token cost than 6 Astra — therealjerseytom · 2026-09-30
- Quantized GLM-5.3 fits on 8x RTX Pro 6000, peaks at 486 decode tok/s — TheZachMueller · 2026-09-30