Language models can control their own attention: 52% decoding cost cut on Gemma 4 31B

jm_alexia · x · 2026-09-04

A new Hugging Face paper page, "Language Models Can Control Their Own Attention," shows models steering their own attention. Zero-shot evaluation on Gemma 4 31B reports a 52% reduction in global attention cost during decoding across 15 long-context benchmarks, with only a 1.52pp accuracy drop. Comments debate whether this rivals chain-of-thought and criticize anthropomorphizing.

Related event: Paper Lets Language Models Control Their Own Attention, Cutting Decode Cost 52%(3 posts)→

Original post →

More from Research

Research channel →