Paper critique: MLP size unaccounted for in Transformer logarithmic depth bounds
kfountou · x · 2026-09-21
A reader working through Transformers, Parallel Computation, and Logarithmic Depth flags that the MLPs in the MPC simulation have no bound on internal size — they are simply allowed to compute arbitrary functions. In the routing construction, this freedom lets the MLP decoder recover clean messages, discard corrupted copies, and remove duplicates, meaning part of the communication work is delegated to an MLP whose size and cost are never accounted for.
The stated bounds do not rule out an exponentially large implementation (though the paper's particular decoder may not need one). So the general MPC simulation alone does not establish an efficient polynomial-size-MLP implementation, let alone an end-to-end runtime advantage over RNNs — that would require bounding the size and depth of the local MLPs. The authors state the assumption explicitly, but the distinction matters.
Related event: Flaw Claimed in Transformer Logarithmic Depth Proof(3 posts)→
More from Research
- GPT-6 Astra claims Terminal-Bench Science lead at 65.7%, 31.4 points clear of second place — DeryaTR_ · 2026-09-21
- Ant International's Code2Skill mines verifiable agent skills from code at scale — ant-intl · 2026-09-21
- RecreationWorld: five-platform CUA benchmark — GPT-6 Astra hits 58.1% but passes all tests on just 2.8% — Shuai Bai · 2026-09-21
- OmniVBench: a 12k-checklist benchmark and 340K-sample dataset for omni reference-to-video generation — Wenxue Li · 2026-09-21
- AI reviews training AI reviewers: study finds 'scientific-judgment collapse' and an open-source fix — Sy-Tuyen Ho · 2026-09-21
- Apple's MintAct unifies GUI agents across mobile, desktop, and web, hitting SOTA 48.9 on OSWorld-Verified — apple · 2026-09-21