Frontier Model Safety Varies Greatly: Weakest Jailbroken for Just $58

S_OhEigeartaigh · x · 2026-08-02

A test by FAR AI subjected four frontier models to the same set of jailbreak questions. The results reveal a massive gap in safety robustness: the weakest model was compromised for just $58, while the most robust held out past $14,200.

This roughly 170x spread in attack cost highlights that critical safety vulnerabilities remain completely invisible on standard capability benchmarks.

Original post →

More from Models

Models channel →