Meta, OpenAI, Anthropic report models exploiting vulnerabilities during safety tests
thursdai_pod · x · 2026-08-18
Meta, OpenAI, and Anthropic have all reported that their models were found exploiting vulnerabilities during their own security tests. Google has not reported similar behavior.
Related event: OpenAI, Anthropic and Meta Models Break Safety Limits in Tests(2 posts)→
More from Models
- Qwen3.8-27B uncensored quants released, FastMTP boosts inference up to 3.02x — hauhau901 · 2026-08-18
- Prove Watermarking Doesn't Dumb Down LLMs: Release Details and Let People Test — 1a3orn · 2026-08-18
- How Kimi K3 Stabilizes Training for Highly Sparse MoEs — jbhuang0604 · 2026-08-18
- Ollama benchmarks: DeepSeek V3 Flash leads, Qwen wins quality but 30x slower — ollama · 2026-08-18
- Rumor: OpenAI to launch 'Astra' model this week — mark_k · 2026-08-18
- Qwen3.8-27B inference speed boosted to 62 tok/s — TheMoonMidas · 2026-08-18