Linear attention suffers poor recall on million-token DNA sequences

No-Coffee-8227 · reddit · 2026-08-16

A researcher modeling DNA sequences (up to 1M tokens) found that linear attention, while efficient, fails at long-range recall, achieving only 25% on Needle-in-a-Haystack tests (random chance). This issue persists even with established models like HyenaDNA. Performance drops as context length increases, despite decent results at 16K context. The author seeks architectural solutions that avoid expensive softmax or external memory.

Related event: Linear Attention Struggles with Recall in Long DNA Sequences(2 posts)→

Original post →

More from Research

Research channel →