vLLM gains RAM offloading: DeepSeek-V4-Flash-Vision-Exp runs on 4x R9700 locally

sloptimizer · reddit · 2026-09-30

Reddit user sloptimizer reports that thanks to tcclaviger's contribution, vLLM now supports RAM offloading, making frontier models far more accessible on local setups. The author ran the original DeepSeek-V4-Flash-Vision-Exp on four AMD R9700 GPUs.

The post includes a full working podman command with notable settings:

Original post →

More from Infra

Infra channel →