Packed-encoders claims 5× faster training by eliminating padding FLOPs
antoine_chaffin · x · 2026-07-28
- The author built packed-encoders, a collection of shape-agnostic CuteDSL and Triton kernels that run on compacted sequences so no FLOPs are wasted on padding.
- Claimed gains: 5× faster training, 2× faster indexing, and 1.2× faster inference.
- The demo code shows how the package wraps a pretrained model and installs a fused forward pass in place, indicating a practical implementation rather than a conceptual proposal.
More from Infra
- Bittensor gets an MCP-style front door for agents to route storage, inference and training — markjeffrey · 2026-07-28
- Moonshot open-sources MoonEP, a balanced MoE communication layer for GPUs and PPUs — deliprao · 2026-07-28
- Post says China is at least a decade away from chips hyperscalers would buy — inductionheads · 2026-07-28
- B200 fine-tuning benchmark shows up to 5.8× speedup and 50% less VRAM — antoine_chaffin · 2026-07-28
- Gemma 4 is benchmarked locally on a 48GB Mac with MLX, llama.cpp and Java 25 — rseroter · 2026-07-28
- Bittensor subnet expansion is pitched as a cheaper AI infrastructure path for companies — markjeffrey · 2026-07-28