Dynamic Expert Migration: Adaptive Optimization from VRAM to RAM

antirez · x · 2026-08-17

This implementation features adaptive logic that migrates routed experts from VRAM to RAM based on past token generation usage, ensuring the hot set remains in VRAM. The author plans to optimize further during a limited two-week access window to reduce slope as context increases.

Related event: Redis Author Optimizes DeepSeek V4 to 45 t/s(3 posts)→

Original post →

More from Infra

Infra channel →