Medusa from training to inference: a two-part guide to multi-token prediction acceleration
No_Progress_5399 · reddit · 2026-08-31
The author wrote a two-part guide explaining Medusa speculative decoding from first principles:
Part 1 (training):
- Medusa heads and the Medusa-1 / Medusa-2 training regimes
- The shifted loss design
- Comparisons with MTP (multi-token prediction)
Part 2 (inference):
- Top-K candidate trees and tree attention
- Verification and rejection sampling
- Greedy versus typical acceptance strategies
The author invites feedback from anyone who has implemented or benchmarked Medusa-style decoding.
- Part 1
- Part 2
More from Infra
- Don't Use LLMs to Quickly Build Multi-Tenant Databases: High Maintenance Cost — kylegawley · 2026-08-31
- ClusterMAX team finds serious security holes in billion-dollar neoclouds — AccBalanced · 2026-08-31
- vLLM: High-Throughput, Memory-Efficient LLM Inference and Serving — goyalshaliniuk · 2026-08-31
- OpenAI reportedly buying tens of thousands of Macs for RL training — SumitGup · 2026-08-31
- Data center reporting should learn from telegraph cable narratives — jwt0625 · 2026-08-31
- New project gpu_api introduces minimal API design, optimizing graphics rendering experience — Michael_Moroz_ · 2026-08-31