Frontier Model Security: Two Models Broken for Under $300
OwainEvans_UK · x · 2026-08-10
In a recent podcast interview, a researcher from FAR AI discussed the current landscape of AI security, specifically focusing on jailbreak protections.
- Test Results: During security testing of four frontier models, two were successfully broken for under $300.
- Vendor Discrepancies: AI companies show vastly different attitudes towards high-severity vulnerabilities. When presented with the same critical universal jailbreak, one developer stated it was "not on our roadmap," while another treated it as a top priority (P0) for immediate resolution.
More from Safety
- "AI Safety" Called a Branding Disaster for Obscuring Core Alignment Issues — jd_pressman · 2026-08-10
- AI compute becomes strategic as tech giants pledge to build their own power infrastructure — bittingthembits · 2026-08-10
- Tsinghua & Cambridge Framework Predicts AI Loss of Control with 84% Accuracy — jiqizhixin · 2026-08-10
- Opinion: High Inference Costs Are the Only Barrier to AI Worms — shlomifruchter · 2026-08-10
- A Mechanistic Explanation of Prompt Injection and Why Roles Matter — katxwoods · 2026-08-10
- Over Half of DEFCON CTF Hackers Now Using AI Coding Assistants — dyn___ · 2026-08-10