A Reddit post says SAOD could shrink a 744B model from 1.5TB to under 100GB
pmttyji · reddit · 2026-07-23
A Reddit post highlights Session-Adaptive Orthogonal Distillation (SAOD), claiming it can compress a 744B model from 1.5 TB down to under 100 GB.
The attached image shows a reproduction-style writeup with plots comparing GRPO, RLSD, and related training behavior. It reports that entropy collapses under GRPO while RLSD stays stable, that the credit-clip ratio lands in the paper’s 3%–6% range, that leakage stays at zero across runs, and that a decay schedule choice meaningfully affects performance. The poster suggests the method could make very large MoE models much more feasible on modest VRAM.
More from Research
- Patch Policy beats a fine-tuned 7B VLA by 18% with 0.7% of the parameters — ylecun · 2026-07-23
- A year-built personal agent was finally beaten by a one-day-old competitor — Antony_Richards · 2026-07-23
- Inkling scores 836 Elo on AA-Briefcase, trailing top open-weight models — ArtificialAnlys · 2026-07-23
- Robotics paper says VLA and world models are not enough for grounded supervision — hbouammar · 2026-07-23
- Google Research: Towards a Quantum Computer That Learns From Its Errors — donutloop · 2026-07-23
- AI could compress decades of biomedical research into days, says Derya Unutmaz — DeryaTR_ · 2026-07-23