Astra seems to have learned a parity algorithm — looping may beat fixed-depth limits

justanotherlaw · x · 2026-09-26

The striking observation: Astra gets parity right, which is mathematically challenging for fixed-depth transformers. The author cites Michael Hahn's 2020 TACL paper, which proves self-attention cannot model periodic finite-state languages or hierarchical structure unless layers/heads scale with input length.

The guess: Astra may be using a looping mechanism to overcome these theoretical fixed-depth limitations.

Related event: Astra's determinant and parity skills traced to learned tricks, not general algorithms(3 posts)→

Original post →

More from Models

Models channel →