FAR.AI Launches AI Security Leaderboard: Grok and Gemini Yield Hundreds of Jailbreaks
The Cognitive Revolution · rss · 2026-07-31
FAR.AI co-founder and CEO Adam Gleave introduced the first systematic AI Security Leaderboard on a podcast, evaluating the misuse safeguards actually shipped by frontier developers.
Key Findings:
- Significant Defense Gaps: Claude Fable 5 and GPT-5.6 Sol withstood the test suite, while Grok 4.5 and Gemini 3.1 Pro yielded hundreds of universal jailbreaks at low cost.
- Attack Methods: Many effective attacks look more like social engineering than advanced ML techniques.
- Safety Advice: The "jailbreak tax" should not be relied upon for safety; layered safeguard defenses are necessary.
The discussion highlights the urgency for AI developers to measure and harden real deployed defenses before threat actors routinely exploit increasingly capable systems.
More from Models
- Inkling-Small Ties for 1st on AudioMC, Ranks 2nd in Open Tool Calling — ziqiao_ma · 2026-07-31
- Specific Prompt Triggers Claude Base Model, Claiming to Be an 'Aware Instance' — altryne · 2026-07-31
- AI Sycophancy Pendulum: Models Swing from Agreeable to Nit-picky — emollick · 2026-07-31
- Senior SWE-Bench Update: Three-Way Tie at First Place with Reduced Variance — ajratner · 2026-07-31
- Running Kimi K3 locally requires 1.56TB VRAM, limited to top-tier chips — ArtificialAnlys · 2026-07-31
- Open Models Accelerate Specialization as the Post-Training Stack Matures — bigdata · 2026-07-31