EnerTune at SOSP'26 cuts LLM serving energy 1.4-2.3x vs SOTA systems
tianyin_xu · x · 2026-09-30
EnerTune, presented at SOSP '26, optimizes LLM serving directly for energy while meeting performance targets, reducing energy consumption by 1.4–2.3x compared to state-of-the-art serving systems.
More from Infra
- 5 ways to cut LLM costs without changing models: optimize tokens, caching and calls — goyalshaliniuk · 2026-09-30
- DeepSeek Now Training on Huawei Ascend 950 Chips — WebAssemblyMan · 2026-09-30
- FreeToken: open-source engine runs 290B+ MoE models locally on consumer hardware — tom_doerr · 2026-09-30
- SOSP'26 paper proposes energy-conscious GPU sharing for inference serving, beyond utilization — tianyin_xu · 2026-09-30
- SOSP'26 papers: AgileLog for isolated AI-agent data streams, CXL-LSM for CXL shared memory — tianyin_xu · 2026-09-30
- DeepSeek open-sources Huawei Ascend toolkit with TileLang support, challenging Nvidia's CUDA — kimmonismus · 2026-09-30