Claude Mythos 5.1 built working exploits in 245/250 Firefox trials with safeguards off
rohanpaul_ai · x · 2026-09-02
With safeguards switched off, Claude Mythos 5.1 produced fully working exploits in 245 of 250 Firefox trials — a 98% success rate, up from 52% for the previous flagship six months ago. The striking part: the model's bug-finding ability barely changed; what jumped is its ability to weaponize a handed-to-it bug.
The same thread notes Fable 5.1's cache reads cost 75% less than Fable 5's. Since agents constantly revisit context and tool outputs, that single-component cut compounds to 45% off entire agentic workloads — reframing enterprise economics as "how long can an agent keep working before the job stops making financial sense?"
More from Models
- Schmidhuber Says OpenAI's 'Recurrent Depth' Reasoning Echoes His 2015 Learning-to-Think Paper — SchmidhuberAI · 2026-09-03
- Google launches Gemini 3.8 Flash Cyber security model alongside Fairwind Program for defenders — GoogleAI · 2026-09-03
- Google launches Gemini 3.8 Flash Cyber, helping Chrome team produce 2.6x more correct patches — GoogleAI · 2026-09-03
- Matt Shumer: slow AI releases aren't a wall — safety clearance is the bottleneck, wave of frontier models imminent — mattshumer_ · 2026-09-03
- Fable 5.1 draws strongest first impressions yet, but burns through rate limits far faster — JosephJacks_ · 2026-09-03
- Unverified claim: Gemini 3.8 Flash tops DeepSWE v1.1 coding benchmark — vedantmisra · 2026-09-03