Opus 5.5 Tops SimpleBench at 88.4%, Xiaomi Open-Sources MiMo-V2.6-Pro Near Frontier
Latent Space · rss · 2026-09-25
Latent Space's weekly AINews roundup: Claude Opus 5.5 leads SimpleBench at 88.4% and is Anthropic's best vision model at 60% lower cost than Fable 5.1; reasoning effort peaks at xhigh (62% on Terminal-Bench-Science) then drops at max. Gemini 3.8 Flash scores 89.2% on ARC-AGI v2 but only 10.4% on v3 standard harness. Xiaomi released MIT-licensed, omni-modal MiMo-V2.6-Pro with 1M context scoring 46 on the AA index (vs 47 for GPT-5.6 Sol) at $0.13/task, open-sourcing its RL code.
TypeSafe's Jev decision model costs $0.044 per 1K judgments—277× cheaper than GPT-6—with a cascade keeping 99% accuracy at 57% cost; the company is reportedly raising $1B+ at a $10B+ valuation. Perplexity's Rust retrieval engine Photon cut p99 from 800ms to 65ms and is now Shopify's main search API. Also: LangChain Deep Agents 0.8, Google flying TPUs in orbit (Project Suncatcher), Liquid AI DSpark 3.13× decode speedup, and GLM-5.3 at 469 tok/s on 8× MI355X.
More from Infra
- Musk details xAI's GPU empire: Colossus 2 hits 110k GB200s with 880k GB300s on the way — ivan_bezdomny · 2026-09-25
- IEEE plenary talk: micro-optimizations across the full stack, from silicon to models — fooobar · 2026-09-25
- Dev builds local AI GTM workflow, argues the next platform entry point is hardware-bound — dotey · 2026-09-25
- US Faces Memory Chip Conundrum as AI-Critical Prices Skyrocket, WSJ Reports — pstAsiatech · 2026-09-25
- Qwen Flash Next IQ4_XS beats 27B FP8 on MMLU-Pro, GPQA and GSM8K in community eval — smallDeltaBigEffect · 2026-09-25
- High-BW TFLN Modulator Demo, but the Dual-Band Grating Coupler Steals the Show — jwt0625 · 2026-09-25