Anthropic says it solved prompt injection, calls OpenAI's new model on par with Gemini Flash
Signalman23 · x · 2026-09-09
Anthropic's Boris Cherny reports that OpenAI's new model scores roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk — a comparison observers called pointed.
Key points:
- Cherny claims Anthropic effectively solved prompt injection for Claude models about two months ago
- He argues publicly evaluating and naming other labs is a great way to push them toward more aligned training, and Anthropic will keep doing it
- He stresses prompt injection remains a serious risk for any model and calls for industry-wide investment in resistance training
A rare public safety rivalry between leading labs.
More from Models
- OpenAI claims huge Navier-Stokes math discovery, academics cry foul — wiredmagazine · 2026-09-09
- OpenAI used a model 'significantly more capable' than Astra for Navier-Stokes, per Axios — 141_1337 · 2026-09-09
- Dwarkesh on MagicAILabs' 50x compute claim: RSI may be less compute-bottlenecked than we think — AccBalanced · 2026-09-09
- Rumor: OpenAI may have cracked Navier-Stokes, with human mathematicians laying years of groundwork — LucaAmb · 2026-09-09
- '50% of open problems just solved' — commentator marvels at frontier model progress — BorisMPower · 2026-09-09
- Cheap models via OpenRouter fall apart in agentic harnesses: GLM and DeepSeek can't match Claude — scottyLogJobs · 2026-09-09