Ablation discussion: query head count barely matters early, sparse routing must be learned
stochasticchasm · x · 2026-08-28
A technical self-reply analyzing two interesting ablations. First, two techniques not usually compared — layer-wise vs sequence-wise indexer sharing — could plausibly be composed, and the author wants a composed ablation. Second, query head count shows almost no impact in the first stage (single vs 4 heads, despite far more FLOPs) but becomes noticeable after stage 2; the author guesses the network must first be trained to route retrieval through sparse blocks, explaining the delayed effect.
More from Research
- Emergent test-time communication proposed as new scaling axis — DimitrisPapail · 2026-08-28
- Benchmark: AI models underperform simple greedy algorithms in retail simulation — ycombinator · 2026-08-28
- Ex-OpenAI Staff: ARC-AGI Pushes False Narrative, Models Capable but Memory Constrained — inductionheads · 2026-08-28
- DiffusionOPSD Cuts Diffusion Model Training Compute by 63% — burny_tech · 2026-08-28
- Paper accepted to EMNLP analyzes geometry of low-resource language LLM representations — davlanade · 2026-08-28
- Google and Peking University introduce PaperBanana for automated scientific figure generation — burkov · 2026-08-28