Agnes 2.5 Pro Beta Uses 50k Output Tokens; Score Rise Driven by Abstention
ArtificialAnlys · x · 2026-08-28
Agnes 2.5 Pro Beta consumes 50k output tokens per task on the Artificial Analysis Intelligence Index, more than double its predecessor Alpha (24k) and exceeding Qwen3.8 27B and GLM-5.3. While its AA-Omniscience score improved from -25 to -11, this is attributed to increased abstention rather than accuracy. It attempted only 45% of questions (vs 94% for Alpha), with accuracy dropping from 33% to 17% despite a lower hallucination rate.
More from Models
- LLM biases differ from human biases, offering utility in reducing bias — JeffLadish · 2026-08-28
- GLM-5.3-Flash hits 270 tok/s: 10% higher quality than 5.2 at one-tenth the cost — Yuchenj_UW · 2026-08-28
- Ornith-1.5-35B-A3B runs agentic coding at 32 tok/s on an 8GB RTX 3070 laptop — Elemental_Particle · 2026-08-28
- User Test: Qwen3.8-Flash-Next Fails at Multi-Turn Conversations — QuixiAI · 2026-08-28
- Alibaba CEO Eddie Wu named to TIME100 AI list; Qwen top open-source model — Aiden_Tech_Ai · 2026-08-28
- Users face 'Forbidden' errors when downloading MiniMax H3 — musashiasano · 2026-08-28