Jenga boosts LLM serving GPU memory utilization by up to 79.6% on vLLM
AccBalanced · x · 2026-08-04
Jenga is a memory-management framework for serving heterogeneous LLMs more efficiently.
- The paper argues that modern LLMs mix different embedding sizes, attention patterns, and access patterns, making traditional allocators inefficient.
- Jenga uses a two-level allocator and LCM-based sizing to reduce fragmentation, while exposing APIs for layer-specific caching and eviction policies.
- Implemented on vLLM, it improves GPU memory utilization by up to 79.6% and is designed to increase batch size and lower inference cost for mixed-model serving workloads.
More from Infra
- Reply says token prices have been falling since late May despite strong demand — GaryMarcus · 2026-08-04
- MiniMax H3 adds open weights, stereo audio and 15-second 2K video generation — petrusenko_max · 2026-08-04
- Qwen 3.8 Max lands on Vercel AI Gateway with low-run cost and mid-pack security scores — evilrabbit_ · 2026-08-04
- How to size a local RAG stack for 50 users, from OCR to reranking — InternationalGap3698 · 2026-08-04
- BlackRock closes $12.5B bond deal for Meta-backed data center — Beth_Kindig · 2026-08-04
- Enterprise AI budgets are already stale as 2027 planning opens — sanjaykalra · 2026-08-04