Sakana AI and Nvidia Unveil TwELL Sparse LLM Solution
SakanaAILabs · x · 2026-07-04
Sakana AI and Nvidia have introduced the TwELL sparse packing format. It adapts to optimized chunked GPU tasks and features custom CUDA kernels to accelerate LLM training and inference. Training and benchmarking on billion-parameter sparse models show a speedup of over 20%, alongside even greater reductions in peak memory and power consumption. The paper will be presented at ICML 2026.
More from Infra
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Vera CPU is speeding up next-gen CPU and GPU design cycles — nordicinst · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27