Final vllm-radiance Build Enables Multi-Agent on One R9700 GPU

KriptacMessage · reddit · 2026-10-11

Reddit user zzpanic shipped the final revision of their custom vllm-radiance build for running Qwen models on a single R9700 GPU, alongside StillDeadcode's new radiance inference engine. Key additions: startup caching for vLLM to skip recompilation on restarts, and a working kv-offload feature that spills KV cache to a VRAM-sysram-disk with memory thinning, making long-context multi-agent work possible on one card. Author says only single-digit-percent speedups remain and has open-sourced the launchers on GitHub.

Original post →

More from coding & agent

coding & agent channel →