State space models bring recurrence back with linear cost and transformer-era tricks
anshulkundaje · x · 2026-10-10
A crisp technical observation: state space models (SSMs) bring recurrence back to sequence modeling while keeping linear cost in sequence length, incorporating transformer-era insights:
- Parallel computation during training, like transformers;
- Multi-head blocks;
- Data-dependent parameters that let each position choose what to write into the state and what to forget — the key differentiator of modern SSMs.
More from Models
- Delip Rao calls out new Qwen3.5-9B-based model for benchmarking latency but not accuracy vs Jev — deliprao · 2026-10-10
- Gemini 4 Argon Launches: 77.9% on DeepSWE v1.1, Beating Claude Opus 5.5's 74.2% — dl_weekly · 2026-10-10
- Polymarket puts 28% odds on Anthropic pausing AI training this month — Polymarket · 2026-10-10
- Netlify livestream blind-tests new Anthropic, OpenAI and Mistral models on real tasks — thisiskp_ · 2026-10-10
- Google reportedly testing Gemini 4 "Carbon" internally; staff say coding "feels like Opus 5.5" — gaganghotra_ · 2026-10-10
- Max reasoning tier costs way more but scores worse on Terminal Bench 4, dev claims — weswinder · 2026-10-10