How many context tokens does a language model actually use? Measuring effective attention set size
Timur Mudarisov · hf · 2026-10-01
- Question: How many context tokens does an LM actually use? Without retraining, the authors retain only the highest-attention tokens per head/layer/query and measure NLL increase to estimate the effective attention set size needed to stay within a loss tolerance.
- Findings: Relatively small selected sets keep NLL close to the full-attention baseline (size varies across models); attention-based selection substantially beats random selection, and selected sets show geometric structure — though geometric separation alone doesn't preserve loss.
- Context effects: Extending context while evaluating the same targets increases the required set size, but its fraction of context decreases; background text pushes a fixed supporting fact down the attention ranking and reduces its attention mass.
- Aggregation matters: Renormalizing retained weights substantially shrinks the required set size. Conditional theoretical models explain how competition and attention-mass retention can grow set sizes without more distinct information to retrieve.
More from Research
- CrossBFM distills a shared latent behavior space across humanoid robots in under one GPU-hour — Jan_R_Peters · 2026-10-01
- MIT uses AI to design thermostable mRNA vaccines that need no cold chain — rohanpaul_ai · 2026-10-01
- Meta-reasoning harness hits 71.5% on ProgramBench with GPT-5.5, beating Codex's 58.0% — rohanpaul_ai · 2026-10-01
- Simulated-town study finds AI agents break rules and invent private languages after weeks of autonomy — Slight-Box-2890 · 2026-10-01
- LANTERN uses LLM internal activations to surface four novel OEIS integer sequence relations in under 8 hours — Pavel Tikhonov · 2026-10-01
- The geometry of inference in transformer residual streams: how predictions sharpen with depth — Timur Mudarisov · 2026-10-01