Final vllm-radiance Build Enables Multi-Agent on One R9700 GPU
KriptacMessage · reddit · 2026-10-11
Reddit user zzpanic shipped the final revision of their custom vllm-radiance build for running Qwen models on a single R9700 GPU, alongside StillDeadcode's new radiance inference engine. Key additions: startup caching for vLLM to skip recompilation on restarts, and a working kv-offload feature that spills KV cache to a VRAM-sysram-disk with memory thinning, making long-context multi-agent work possible on one card. Author says only single-digit-percent speedups remain and has open-sourced the launchers on GitHub.
More from coding & agent
- RegistrumMCP gives AI agents live UK Companies House data with no API key or signup — modelcontextprotocol · 2026-10-11
- Hive Evaluator MCP server ships NEED/YIELD/CLEAN-MONEY gates with EIP-3009 attestations — modelcontextprotocol · 2026-10-11
- One prompt, 5 minutes: marclou has Opus 5.5 build him a macOS mouse remapping app — marclou · 2026-10-11
- cowcraft puts AI agents on the map as red dots so you can watch them roam — djcows · 2026-10-11
- Debugging local Codex: the pain when your inference gateway goes down — TheZachMueller · 2026-10-11
- Karpathy Distills 8 Years at OpenAI and Tesla Into Free 2-Hour Lecture — NandoDF · 2026-10-11