Training an EAGLE-3 Speculative Decoding Drafter for Gemma-3-27B on a Single RTX 5090

max_paperclips · x · 2026-08-11

Developer witcheer shared a detailed breakdown of training an EAGLE-3 speculative decoding drafter for Gemma-3-27B from scratch. The training was completed in just four nights on a single RTX 5090 GPU.

The 716M-parameter drafter head uses a single-layer Llama architecture and is optimized for the sglang framework. Benchmarks show significant inference speedups with zero loss in output quality:

The author also evaluated various tree-based speculative configurations, noting that tree-3-4-8 yields the best net speedup. The model weights have been open-sourced on Hugging Face.

Original post →

More from Infra

Infra channel →