Opus 5.5 tops SimpleBench as code models emerge as explainer-video engines; MiMo-V2.6-Pro open-sourced
Latent Space · rss · 2026-09-29
Latent Space AINews (9/24–9/25) covers a frontier model wave and much more:
- Models: Opus 5.5 leads SimpleBench at 88.4%; reasoning effort analysis shows 24%→62%→59% from low→xhigh→max on Terminal-Bench-Science. GPT-6 Astra reportedly beat NetHack. Gemini 3.8 Flash scores 89.2% on ARC-AGI v2. Xiaomi open-sourced omni-modal MiMo-V2.6-Pro under MIT (AA index 46, $0.13/task) with RL code.
- "System One" decision models: TypeSafe's Jev delivers typed decisions at $0.044/1K judgments (277× cheaper than GPT-6), proved 140 theorems for under $1; alternatives CLM, GLiNER2.5-Decide, Tev1.
- Infra: LangChain's Managed Deep Agents 0.8; Perplexity's Rust retrieval engine Photon cut p99 from 800ms to 65ms, now Shopify's main search API; Liquid AI DSpark gives 3.13× decode speedup; GLM-5.3 hits 469 tok/s on 8× MI355X; Google is flying TPUs in orbit.
- Research: Harness-Zero distills agent harnesses into models (23.3%→44.3% without a harness); XYEval shows one misleading user hint cuts scores up to 46.7%; a NeurIPS paper shows a single MLP neuron suppression bypasses refusals across 7 models; Sakana hires Schmidhuber as Chief Scientific Advisor of its RSI Lab.
- World models & media: Odyssey's Agora-2 simulates 20 agents in real time; Opus 5.5/Astra produce videos entirely from code, prompting "who knew you didn't need diffusion" takes.
More from Infra
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x — BAAI · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29