Claude Opus 5.5 tops AA Index at 58, but per-task cost stays flat as output volume jumps 60%
qinzytech · x · 2026-10-07
Claude Opus 5.5 tops the Artificial Analysis Intelligence Index at 58, but token prices falling 20% hasn't lowered per-task cost: at max effort it emits 119k output tokens per task vs 73k for Opus 5, at $4/M input and $20/M output.
- Leads 6 of 10 Index evals; scores 59.6% on Terminal-Bench 4.0, matching GPT-6 Astra at xhigh
- AA-Briefcase Elo 1822, beating Claude Fable 5.1 on analytical quality and presentation
- Four of five effort settings sit on the intelligence-cost frontier
- Still slightly behind on CritPt, AA-LCR and GDP.pdf
More from Models
- Ex-Googler suspects Project Astra voice got quantized and downgraded — joannejang · 2026-10-07
- Running quantized Swift 1.5 in 16GB VRAM at 50 tps: can lower quants help agentic coding? — royalflash417 · 2026-10-07
- OpenAI Reportedly Makes Serious Breakdown Toward Riemann Hypothesis — ipeirotis · 2026-10-07
- Jev-as-a-Judge: New Model Boosts LLM Judge Reliability for Agent Evals — omarsar0 · 2026-10-07
- X rolls out @bot tagging: reply to any post to save to Notion, set reminders via Grok — tetsuoai · 2026-10-07
- Researcher: OpenAI's math results are 'bonkers' — labs should push it to medicine next — Afinetheorem · 2026-10-07