SOSP'26 paper proposes energy-conscious GPU sharing for inference serving, beyond utilization
tianyin_xu · x · 2026-09-30
At SOSP'26, NeerajaJY's team is presenting "Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving", arguing that GPU sharing for inference serving should go beyond utilization metrics and take energy consumption into account. Paper and thread via the original post.
More from Infra
- 5 ways to cut LLM costs without changing models: optimize tokens, caching and calls — goyalshaliniuk · 2026-09-30
- DeepSeek Now Training on Huawei Ascend 950 Chips — WebAssemblyMan · 2026-09-30
- FreeToken: open-source engine runs 290B+ MoE models locally on consumer hardware — tom_doerr · 2026-09-30
- EnerTune at SOSP'26 cuts LLM serving energy 1.4-2.3x vs SOTA systems — tianyin_xu · 2026-09-30
- SOSP'26 papers: AgileLog for isolated AI-agent data streams, CXL-LSM for CXL shared memory — tianyin_xu · 2026-09-30
- DeepSeek open-sources Huawei Ascend toolkit with TileLang support, challenging Nvidia's CUDA — kimmonismus · 2026-09-30