GeoPair: Training-Free Cross-Layer Factorization Hits SOTA in Transformer Compression
MTSAIR · hf · 2026-09-24
MTSAIR introduces GeoPair, a principled training-free framework for post-training transformer compression. Instead of optimizing layers in isolation or using heuristic grouping, it sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations that preserve each layer's calibration geometry.
Combined with structured sparsity, this yields efficient weight decompositions that consistently outperform independent structured decomposition and alternative pairwise factorizations across architectures, scales, and modalities, replacing heuristic engineering with a convergent optimization-driven pipeline.
More from Infra
- Ban ultra-high-bandwidth interconnects, not GPUs, to stop large-scale AI training — davidmanheim · 2026-09-24
- 10 open-source GitHub repos for building your own AI inference stack — Shruti_0810 · 2026-09-24
- Qualcomm commits to official Linux support for Snapdragon X2, Ubuntu certification by H1 2027 — carrycooldude · 2026-09-24
- Alibaba bets across the full stack: 5-10T Qwen models, 500K-card clusters, 20GW cloud by 2032 — Div_pradeep · 2026-09-24
- Chose rustpython-parser over tree-sitter after measuring both on real LLM output — SprayPuzzleheaded533 · 2026-09-24
- More compute made one agent 13x faster, another just 4%: agents' bottlenecks are task-dependent — alex_verem · 2026-09-24