Kubernetes DRA Reaches GA: Native GPU Scheduling and Slicing
sloppenheimer · x · 2026-08-13
Kubernetes' Dynamic Resource Allocation (DRA) reached general availability in v1.34 and is enabled by default since v1.35, natively supporting GPU scheduling.
A maintainer from CNCF incubating project HAMi breaks down DRA's impact on existing GPU sharing solutions. Previously, due to the narrow device plugin API, HAMi built a complex pipeline to achieve GPU memory and compute slicing. Now, DRA's consumable capacity feature allows Pods to request fractions of a device directly from the scheduler, absorbing HAMi's resource allocation role.
However, DRA does not make HAMi obsolete. While DRA handles scheduling, HAMi continues to provide hard in-container enforcement at the CUDA-call granularity. HAMi has decoupled its enforcement layer from DRA-based allocation to adapt to the new ecosystem.
More from Infra
- Red Hat's DSpark Speculator Boosts Kimi-K3 Throughput by 3.5x — teortaxesTex · 2026-08-14
- AMD Raising $5 Billion in Debt to Fund AI War Against Nvidia — ns123abc · 2026-08-14
- Merge Gateway Launches Multimodal API for Unified Access to Image, Video, Audio Models — shensi · 2026-08-14
- MasterClass Adopts CoreWeave and W&B Weave to Monitor AI Teaching Agents — wandb · 2026-08-14
- SkyPilot AI Infra Meetup next Tuesday in SF with VAST Data and NVIDIA — skypilot_org · 2026-08-14
- SanDisk Predicts Flash Market to Approach $500B by 2027 Amid AI Boom — zephyr_z9 · 2026-08-13