Expert skeptical of distilling Transformers to RNNs, cites architectural gap
ChengleiSi · x · 2026-09-01
Albert Gu critiqued a paper claiming sliding-window attention beats linear attention post-training, labeling it misleading 'retrofitting' rather than post-training. He expressed bearishness on distilling Transformers to recurrent models due to architectural differences, arguing new architectures should be trained from scratch.
More from Research
- EMNLP 2026 Budapest Tutorial on Multilingual and Multicultural LLMs Announced — cocoweixu · 2026-09-21
- Xiaomi's CodeMidas turns source code into RL environments: 5,545 tasks, 23 languages — teortaxesTex · 2026-09-21
- Delip Rao responds to COLM paper code similarity claim: repo and tech reports public since Feb — deliprao · 2026-09-21
- Nina Miolane argues AI safety needs 'engineering laws': predictive math for intelligence — ninamiolane · 2026-09-21
- Safety follows understanding: Nina Miolane calls for predictive laws of intelligence — ninamiolane · 2026-09-21
- UC Berkeley launches Science of Intelligence Institute; Nina Miolane to give inaugural lecture Oct 7 — ninamiolane · 2026-09-21