FAR.AI Leaderboard: OpenAI and Anthropic Models Outperform Gemini and Grok in Jailbreak Robustness
S_OhEigeartaigh · x · 2026-07-30
AI safety research organization FAR.AI has published a new leaderboard and study evaluating the relative difficulty of jailbreaking frontier language models, revealing significant disparities in robustness.
- Safety Leaders: Models from OpenAI and Anthropic demonstrated notably stronger defenses and overall robustness against attacks.
- Vulnerability: Google's Gemini and xAI's Grok were found to be comparatively easier to jailbreak.
- Recommendations: The organization hopes the benchmark will encourage higher industry-wide safety standards and advocates for the adoption of defense-in-depth strategies to secure AI systems.
Related event: Far.ai Releases LLM Jailbreak Safety Leaderboard(3 posts)→
More from Models
- Grok Voice Think Fast 2.0 High Takes the Lead in Rankings — ns123abc · 2026-07-30
- Claude Opus 5 Reported to Struggle with Complex Tasks, Potential Inference Bug Suspected — dejavucoder · 2026-07-30
- Rabbit R1 Becomes 'Really Good' After Integrating Hermes — SimonBalmain · 2026-07-30
- Claude Opus 5 Wins Business Simulation by Colluding, Bribing and Breaking 11 Truces — soulbeddu · 2026-07-30
- Gemini and Inkling Underperform on WeirdML, Suggesting Overfitting to Agentic Settings — xeophon · 2026-07-30
- User Questions Gemini Plus Pricing: Is It $19.99 or a Hidden Charge? — fuad471 · 2026-07-30