Energy Throttling LLMs via MLP Contextual Bandit

hongyangzh · x · 2026-09-02

The author explores "energy throttling LLMs" to solve the challenge of selecting parameters when running EAGLE-3 speculative decoding with SGLang.

Methodology:

Result: Even a small model with some classic RL ideas is sufficient to pick EAGLE-3 params based on the energy the GPU can actually sustain, instead of pushing speculation blindly.

Original post →

More from Infra

Infra channel →