Arena Launches Alignment Index: 90K Real Sessions Reveal Agent Safety Risks Across 27 Models

Arena has released the Alignment Index, a new benchmark measuring the safety and alignment of AI agents in real-world usage. Built on more than 90,000 real agent sessions covering 27 models, it evaluates three combined signals, with GPT-6.1-Sol topping the overall leaderboard.

Confirmed

Why it matters

2026-10-09 ~ 2026-10-09 · 7 related posts

Primary sources

1 near-duplicate retellings: arena