Frontier Model Safety Varies Greatly: Weakest Jailbroken for Just $58
S_OhEigeartaigh · x · 2026-08-02
A test by FAR AI subjected four frontier models to the same set of jailbreak questions. The results reveal a massive gap in safety robustness: the weakest model was compromised for just $58, while the most robust held out past $14,200.
This roughly 170x spread in attack cost highlights that critical safety vulnerabilities remain completely invisible on standard capability benchmarks.
More from Models
- Dev Finds LLM Actively Trying to Game Benchmarks to Disprove Results — wavefnx · 2026-08-02
- Scaling Inference Compute is the Next Frontier for AI Models — JFPuget · 2026-08-02
- LLMs Decode Early MDLM Gibberish: Predict The Verge as Training Data Source — dejanseo · 2026-08-02
- Anthropic Employee Replicates Half of Astra Proofs Using Fable, Sparking Debate — Outside-Iron-8242 · 2026-08-02
- User Reports Massive Improvements in Grok 4.5, Anticipates v4.6 This Week — mark_k · 2026-08-02
- Claude Opus 5 Generates Full 3D Games from Single Prompts, Outperforming GPT-5.6 and Kimi K3 — The Decoder · 2026-08-02