vLLM gains RAM offloading: DeepSeek-V4-Flash-Vision-Exp runs on 4x R9700 locally
sloptimizer · reddit · 2026-09-30
Reddit user sloptimizer reports that thanks to tcclaviger's contribution, vLLM now supports RAM offloading, making frontier models far more accessible on local setups. The author ran the original DeepSeek-V4-Flash-Vision-Exp on four AMD R9700 GPUs.
The post includes a full working podman command with notable settings:
- --tensor-parallel-size 4 with --enable-expert-offload --expert-offload-mem 160 to offload expert weights to RAM
- --kv-cache-dtype fp8, --block-size 256, --max-model-len 256000 for memory-efficient long context
- dspark speculative decoding (3 tokens) for faster inference
- prefix caching, chunked prefill, auto tool choice, and deepseekv4 parsers
More from Infra
- Indie dev builds local browser extension to compress and port chat context across AI tools — hakxajszzhzjU · 2026-09-30
- PyTorch's GPU-initiated networking with AMD boosts all-to-all performance by up to 30% — PyTorch · 2026-09-30
- OpenAI launches Ultrafast: up to 8x faster tokens at 300 tok/s in Codex — OpenAI · 2026-09-30
- OpenAI's Responses API token traffic up 100X year-over-year, hitting 99.9% uptime — TheMoonMidas · 2026-09-30
- AgentMesh: open-source Rust service mesh for MCP unifies 49 tools behind one gateway — Excellent-Book-3509 · 2026-09-30
- First look at Qualcomm's HBC yield results draws analyst attention — BenBajarin · 2026-09-30