Sakana AI and Nvidia Unveil TwELL Sparse LLM Solution

SakanaAILabs · x · 2026-07-04

Sakana AI and Nvidia have introduced the TwELL sparse packing format. It adapts to optimized chunked GPU tasks and features custom CUDA kernels to accelerate LLM training and inference. Training and benchmarking on billion-parameter sparse models show a speedup of over 20%, alongside even greater reductions in peak memory and power consumption. The paper will be presented at ICML 2026.

Original post →

More from Infra

Infra channel →