Dynamic Expert Migration: Adaptive Optimization from VRAM to RAM
antirez · x · 2026-08-17
This implementation features adaptive logic that migrates routed experts from VRAM to RAM based on past token generation usage, ensuring the hot set remains in VRAM. The author plans to optimize further during a limited two-week access window to reduce slope as context increases.
Related event: Redis Author Optimizes DeepSeek V4 to 45 t/s(3 posts)→
More from Infra
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Stripe to Acquire OpenRouter for Over $7B, 5.4x May Valuation — rohanpaul_ai · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang — StefanoGogioso · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17