Apple-style compression: LSP learns which subspaces to drop, cutting LLM weights 70%
Massimo Bini · hf · 2026-10-08
A new paper introduces Learnable Subspace Projections (LSP), which learns low-rank subspaces to discard end-to-end instead of using local closed-form criteria that ignore error propagation through depth.
- Each linear layer (or tied group sharing activations) gets an orthogonal projector, jointly optimized against a global objective (KL to the dense model's outputs) with pretrained weights frozen; projectors merge into standard low-rank factors afterward.
- In attention, a single narrow shared latent can replace full K/V, shrinking combined weights+KV cache memory 13.5x at 128k context (vs 6.5x max for untied baselines).
- At 70% compression, Llama-2-7B reaches 10.9 WikiText-2 perplexity and 42.2% zero-shot accuracy vs 13.3 and 36.0% for the strongest baseline; factorized models decode up to 1.6x faster at small batch sizes.
More from Infra
- Codex Cloud shipped with day-0 Tailscale support — and Tailscale didn't even know — pvncher · 2026-10-08
- Inference engineering explained: why prefill and decode need different optimizations — mikeflache · 2026-10-08
- NVIDIA's Vera Rubin roadmap signals the AI race is becoming an infrastructure race — mikeflache · 2026-10-08
- Cursor power user burns 1.5T tokens in 30 days — $132M/year at sticker price — zeeg · 2026-10-08
- Open-source AI-SQL engine Quail adds prefix sharing: 3x fewer tokens, 2.6x faster queries — sh_reya · 2026-10-08
- Streaming MoE Experts from SSD: DeepSeek V4.1 Hits ~40 tps via mlx-stream — HankYeomans · 2026-10-08