SSR: Self-Speculation for Faster Reasoning Models

kastnerkyle · x · 2026-08-25

Researchers from UCLA introduce SSR, a training-free self-speculative decoding method to accelerate models with long Chain-of-Thought (CoT) traces. SSR uses the partial-CoT distribution as a drafter and the full-CoT distribution as a verifier, leveraging semantic and lexical overlap to accept long draft prefixes at once. It also incorporates suffix decoding to recover useful spans beyond the prefix, reducing latency in structured and long-form generation tasks.

Original post →

More from Infra

Infra channel →