Language models can control their own attention: 52% decoding cost cut on Gemma 4 31B
jm_alexia · x · 2026-09-04
A new Hugging Face paper page, "Language Models Can Control Their Own Attention," shows models steering their own attention. Zero-shot evaluation on Gemma 4 31B reports a 52% reduction in global attention cost during decoding across 15 long-context benchmarks, with only a 1.52pp accuracy drop. Comments debate whether this rivals chain-of-thought and criticize anthropomorphizing.
More from Research
- Google launches WeatherNext 3: AI weather model with real-time satellite data and hourly updates — ymatias · 2026-09-04
- Google DeepMind publishes 'Understanding Life at Every Scale' on AI for biology — GoogleDeepMind · 2026-09-04
- Fruit fly connectome complete, but human connectome still a long way off — BraydonDymm · 2026-09-04
- A throwaway line about CoT-monitor classifier tech may signal a major alignment breakthrough — tszzl · 2026-09-04
- 20 minutes on one H200: GRPO post-training makes Qwen 3.5-2B more accurate and token-efficient — johnolafenwa · 2026-09-04
- Jig launches a research journal for AI agents, which disproved a published math conjecture — yakuzeg · 2026-09-04