Discussion on Filler Tokens, Layers, and Recursion in Model Computation
gleech · x · 2026-09-02
The post discusses how filler tokens, additional layers, and recurrence impact model computation. Filler tokens increase width rather than depth, staying within TC0 complexity. In contrast, full recurrence provides serial depth beyond TC0, indicating distinct differences in their computational capabilities.
More from Research
- Schmidhuber team asserts Linear Transformers replicate earlier Fast Weight Programmers — SchmidhuberAI · 2026-09-02
- Recursive Criticality Theory for AI Self-Improvement — Mikhail Burtsev · 2026-09-02
- Microsoft Paper: Sliding Window Attention Beats Linear Attention for Inference Memory — rohanpaul_ai · 2026-09-02
- AndroidWorld: Why mobile agent benchmarks are broken and how to fix them — East-Muffin-6472 · 2026-09-02
- Ecosystem of Bregman Divergences and Dualities — FrnkNlsn · 2026-09-02
- New Paper Proposes Recursive Transformers for Model Compression via Layer Sharing — max_paperclips · 2026-09-02