SSR: Self-Speculation for Faster Reasoning Models
kastnerkyle · x · 2026-08-25
Researchers from UCLA introduce SSR, a training-free self-speculative decoding method to accelerate models with long Chain-of-Thought (CoT) traces. SSR uses the partial-CoT distribution as a drafter and the full-CoT distribution as a verifier, leveraging semantic and lexical overlap to accept long draft prefixes at once. It also incorporates suffix decoding to recover useful spans beyond the prefix, reducing latency in structured and long-form generation tasks.
More from Infra
- Engineering postmortem: False 100% success rates and hidden latency disasters — CupGlass540 · 2026-08-25
- Open Source MiniMax H3 Inference Engine for Mac — QuixiAI · 2026-08-25
- H3 Inference Benchmarks: 8x 3090 Beats M3 Ultra — QuixiAI · 2026-08-25
- AI Drives First US Grid and Blue-Collar Job Expansion in 20 Years — danielrock · 2026-08-25
- Relace cuts Deepseek v3 inference costs by nearly 50% on OpenRouter — stuffyokodraws · 2026-08-25
- Garry Tan: Data Center NIMBYs Are Destroying $1 Trillion a Year in AI Infrastructure — garrytan · 2026-08-25