Training an EAGLE-3 Speculative Decoding Drafter for Gemma-3-27B on a Single RTX 5090
max_paperclips · x · 2026-08-11
Developer witcheer shared a detailed breakdown of training an EAGLE-3 speculative decoding drafter for Gemma-3-27B from scratch. The training was completed in just four nights on a single RTX 5090 GPU.
The 716M-parameter drafter head uses a single-layer Llama architecture and is optimized for the sglang framework. Benchmarks show significant inference speedups with zero loss in output quality:
- Code generation: 1.44x speedup (59.6 to 89.0 tok/s)
- Repetitive text: 1.52x speedup
- Prose: 1.33x speedup
- Chat: 1.23x speedup
The author also evaluated various tree-based speculative configurations, noting that tree-3-4-8 yields the best net speedup. The model weights have been open-sourced on Hugging Face.
More from Infra
- OpenAI Sends Letter to Texas Governor on Responsible AI Infrastructure — ArtificialOther · 2026-08-11
- Optimizing NVFP4 Blockscaled GEMM on RTX Pro 6000 Blackwell — HanGuo97 · 2026-08-11
- CoreWeave Expands into APAC with 360MW Data Centers in Indonesia — Beth_Kindig · 2026-08-11
- OpenAI Believed to Cut Datadog Usage, Impacting Cloud Provider's Guidance — SumitGup · 2026-08-11
- Microsoft Plans to 'Significantly' Increase Production of Next-Gen AI Chips — thoefler · 2026-08-11
- fal Signs 3 Hot GenAI Model Companies, Expands H200 and B300 Capacity — gorkem · 2026-08-11