Dynamic Short Convolutions Improve Transformer Performance

Cohere · youtube · 2026-08-15

Cohere Labs hosts a talk by MIT PhD student Oliver Sieberling on "Dynamic Short Convolutions." This approach uses input-dependent filters instead of static ones, allowing adaptive aggregation of local context to improve the expressivity and performance of Transformer-based language models. The talk also covers efficient implementation using custom Triton kernels.

Original post →

More from Infra

Infra channel →