Long Context's Softmax Denominator Problem: Attention Dilution and Sparse Fixes

CShorten30 · x · 2026-10-02

A Weaviate podcast clip breaks down why long context windows hit a math problem in the softmax denominator: it sums over every token while the numerator for the relevant token stays fixed, flattening attention scores until the right document barely stands out. Two mitigations are discussed: BlockSearch uses length-aware softmax scaling and drops low-scoring documents before attention runs, while REFRAG lets the decoder read pre-computed chunk embeddings and expand only the chunks a lightweight policy deems necessary. The episode also covers the state of sparse attention and synergies with vector databases.

Original post →

More from Research

Research channel →