Dynamic Short Convolutions Improve Transformer Performance
Cohere · youtube · 2026-08-15
Cohere Labs hosts a talk by MIT PhD student Oliver Sieberling on "Dynamic Short Convolutions." This approach uses input-dependent filters instead of static ones, allowing adaptive aggregation of local context to improve the expressivity and performance of Transformer-based language models. The talk also covers efficient implementation using custom Triton kernels.
More from Infra
- Oracle's Stargate AI Data Center Gas Pipeline Delayed to Feb 2027 — Polymarket · 2026-08-15
- Core AI Model Zoo Enables Local LLM/VLM Inference on Apple Devices — amos_gyamfi · 2026-08-15
- NVIDIA CEO Jensen Huang announces AI factory infrastructure as new asset class with BlackRock, Blackstone, and others — ___Mufasaa · 2026-08-15
- Advantech launches AFE-A702 edge device with Thor to sync 8 RealSense cameras — chrismatthieu · 2026-08-15
- Soleio: Compute is sovereignty — soleio · 2026-08-15
- CEO fixes MoE kernel index overflow bug in sonic-moe — rosstaylor90 · 2026-08-15