ICML paper: log-depth transformers suffice via equivalence with parallel computation
kfountou · x · 2026-09-21
- Sanford, Hsu & Telgarsky (ICML 2024) show constant-depth self-attention efficiently simulates — and is simulated by — a constant number of rounds of Massively Parallel Computation (MPC).
- Consequence: logarithmic depth suffices for transformers to solve basic tasks that several other neural sequence models and sub-quadratic approximations cannot solve efficiently, establishing parallelism as the key distinguishing property of transformers.
- kfountou raises a technical question while studying the paper: the MLPs in the MPC simulation are unbounded and allowed to compute arbitrary functions, letting the routing construction recover clean messages, discard corrupted copies and deduplicate — but that freedom makes the construction theoretically murky.
Related event: Flaw Claimed in Transformer Logarithmic Depth Proof(3 posts)→
More from Research
- GPT-6 Astra claims Terminal-Bench Science lead at 65.7%, 31.4 points clear of second place — DeryaTR_ · 2026-09-21
- Ant International's Code2Skill mines verifiable agent skills from code at scale — ant-intl · 2026-09-21
- RecreationWorld: five-platform CUA benchmark — GPT-6 Astra hits 58.1% but passes all tests on just 2.8% — Shuai Bai · 2026-09-21
- OmniVBench: a 12k-checklist benchmark and 340K-sample dataset for omni reference-to-video generation — Wenxue Li · 2026-09-21
- AI reviews training AI reviewers: study finds 'scientific-judgment collapse' and an open-source fix — Sy-Tuyen Ho · 2026-09-21
- Apple's MintAct unifies GUI agents across mobile, desktop, and web, hitting SOTA 48.9 on OSWorld-Verified — apple · 2026-09-21