Gumbel Straight Flow: distilling autoregressive models into one-step flow maps
sedielem · x · 2026-10-02
Gumbel Straight Flow flips the usual distillation setup: can autoregressive (language) models serve as teachers for flow map parallel samplers? The method distills AR models into one-step flow maps, enabling parallel sampling instead of token-by-token generation.
More from Research
- Company Knowledge Bench: real-world retrieval benchmark tests 7 retrievers — CShorten30 · 2026-10-03
- Gacha Decoding diversifies LLM outputs with 11x better sample efficiency — natolambert · 2026-10-03
- New paper sparks debate: intelligence vs parroting is a computational distinction — ctjlewis · 2026-10-03
- RLE-Bench grades coding agents as robot learning engineers across four workflows — RLE-Bench · 2026-10-03
- Distillation study: KL direction and learning rate matter more than on-policy rollouts — CambUni · 2026-10-03
- Stanford revisits 2018 claim that self-driving was 90% done — the last 10% was everything — StanfordHAI · 2026-10-03