Linear attention suffers poor recall on million-token DNA sequences
No-Coffee-8227 · reddit · 2026-08-16
A researcher modeling DNA sequences (up to 1M tokens) found that linear attention, while efficient, fails at long-range recall, achieving only 25% on Needle-in-a-Haystack tests (random chance). This issue persists even with established models like HyenaDNA. Performance drops as context length increases, despite decent results at 16K context. The author seeks architectural solutions that avoid expensive softmax or external memory.
Related event: Linear Attention Struggles with Recall in Long DNA Sequences(2 posts)→
More from Research
- Retriever: A Framework for Asynchronous, Closed-Loop Robot Agents — ZeYanjie · 2026-08-24
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24