Agnes 2.5 Pro Beta Joins Frontier Tier with Major Agentic Gains, Higher Cost
ArtificialAnlys · x · 2026-08-28
Singapore-based lab Agnes AI released Agnes 2.5 Pro Beta, scoring 49 on the Artificial Analysis Intelligence Index—a 9-point jump from the Alpha version. This places it in the frontier-adjacent tier, just behind Gemini 3.5 Flash and GPT-5.6 Luna, and ahead of MiniMax-M3.
The primary driver is a leap in agentic capabilities: the Agentic Index rose from 25 to 44. The τ³-Banking metric nearly tripled, and GDPval-AA v2 Elo score increased significantly against a human baseline. Frontier reasoning benchmarks saw modest gains, with Humanity's Last Exam rising from 34% to 38%.
The model significantly reduced hallucination rates (from 88% to 33%) by adopting a more conservative abstention strategy, attempting only 45% of questions compared to the Alpha's 94%. However, this comes at a cost: output token usage doubled to 50k per task. The model features a 1M token context window, supports text/image input, and is priced at $0.10/$0.30 per 1M input/output tokens.
Related event: Agnes 2.5 Pro Beta Jumps in Rankings at Doubled Token Cost(4 posts)→
More from Models
- Users face 'Forbidden' errors when downloading MiniMax H3 — musashiasano · 2026-08-28
- User Opinion: GPT Sol Would Be the Best Model With Better Front-End Skills — BLUECOW009 · 2026-08-28
- Anima-3.8B Trends on Hugging Face — lylogummy · 2026-08-28
- Evolution strategies beat GRPO on reasoning diversity and Pass@K, paper finds — Yunpeng Ba · 2026-08-28
- Opinion: Qwen's n-gram innovation pops the AI bubble — Acrobatic_Stress1388 · 2026-08-28
- IBM open sources Granite 4.2 reasoning models with full training recipe — krvarshney · 2026-08-28