Expert skeptical of distilling Transformers to RNNs, cites architectural gap

ChengleiSi · x · 2026-09-01

Albert Gu critiqued a paper claiming sliding-window attention beats linear attention post-training, labeling it misleading 'retrofitting' rather than post-training. He expressed bearishness on distilling Transformers to recurrent models due to architectural differences, arguing new architectures should be trained from scratch.

Original post →

More from Research

Research channel →