NAPE Audio Pretraining Achieves SOTA Without Decoders

kastnerkyle · x · 2026-08-24

UmbertoSenpai and LiuXub introduce NAPE, a new self-supervised audio learning method. By causally predicting the next audio patch embedding with stop-grad, NAPE achieves strong and scalable audio learning performance. It reaches SOTA on benchmarks, offers interpretable results, and requires no decoders, teacher-student models, or regularization losses. The method is simple, scalable, and seamlessly integrates native audio pretraining into multimodal LLMs.

Original post →

More from Multimodal

Multimodal channel →