Experiment calibrates 48 attention heads down to 12, swapping the rest for band-diagonal sparse attention
ostrisai · x · 2026-09-18
- Lodestone Rock shares an experiment arguing full attention is overkill and should be used sparingly
- They calibrated krea2 from 48 full-attention heads down to 12, converting the rest to scanline (band-diagonal) attention that is block-sparse by design
- The accompanying gif visualizes the sparse attention pattern
More from Models
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18
- Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans — amplifiedamp · 2026-09-18
- Noam Brown: models may perform their chain of thought; alignment must be solved — infoxiao · 2026-09-18
- Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings — amplifiedamp · 2026-09-18
- Independent eval puts new model Jev at Terra no-think level, roughly on par with Luna-xhigh — tokenbender · 2026-09-18
- Hugging Face adds zero-shot text classification pipeline to its API, no training needed — joeddav · 2026-09-18