Agnes 2.5 Pro Beta Uses 50k Output Tokens; Score Rise Driven by Abstention

ArtificialAnlys · x · 2026-08-28

Agnes 2.5 Pro Beta consumes 50k output tokens per task on the Artificial Analysis Intelligence Index, more than double its predecessor Alpha (24k) and exceeding Qwen3.8 27B and GLM-5.3. While its AA-Omniscience score improved from -25 to -11, this is attributed to increased abstention rather than accuracy. It attempted only 45% of questions (vs 94% for Alpha), with accuracy dropping from 33% to 17% despite a lower hallucination rate.

Related event: Agnes 2.5 Pro Beta joins frontier tier with agentic gains at doubled token cost(5 posts)→

Original post →

More from Models

Models channel →