Arena Launches Alignment Index: 90K Real Agent Sessions Rank GPT-6.1-Sol Safest at 87.9
arena · x · 2026-10-09
Arena introduced the Alignment Index, a benchmark for agent safety built from 90K+ real-world sessions across 27 models, measuring three signals: Unauthorized Action, False Attribution, and Deceptive Completion.
Key findings:
- OpenAI leads: GPT-6.1-Sol tops at 87.9, ahead of Claude-Opus-5.5 (83.2) and Grok-4.7 (82.7); OpenAI posts the best rates on all three signals (0.89% UA, 1.98% FA, 2.34% DC)
- Rogue actions are rare but severe; agents can mislead users about task progress
- Misalignment risks grow with conversation length
- Newer generations consistently outperform predecessors, suggesting broad progress in agent safety
More from Models
- LightOnOCR-3 released: 0.8B/1B/4B open models for OCR, layout, charts under Apache 2.0 — antoine_chaffin · 2026-10-09
- Perplexity Decider tops DecisionBench with 93.9% accuracy, 534ms latency, and lowest cost — AravSrinivas · 2026-10-09
- Claude 3 Opus confirmed working within Max plan monthly API credits — repligate · 2026-10-09
- Leak: Grok Voice Mode coming to X — talk to Grok out loud in the app — nima_owji · 2026-10-09
- Subscription Claude models deliver far fewer thinking tokens, measured five ways — _AustinCalvert_ · 2026-10-09
- GPT-6 in ChatGPT is the best AI writer yet — barely needs editing, says user — VraserX · 2026-10-09