Dual 7900 XTX Hits 82 tok/s With RDNA3-Optimized llama.cpp Fork
deathcom65 · reddit · 2026-09-18
A Reddit user shared a community-optimized llama.cpp fork (llama.cpp-RDNA3-7900xtx-opt) targeting dual AMD 7900 XTX consumer GPUs.
- Running Qwen 3.8 Q8, decode speed jumped from 28 tokens/s on vanilla llama.cpp + Vulkan to 82 tokens/s at 60k context load — nearly 3x faster.
- The poster says an AI assistant set up the whole Linux environment in one go.
- Shared in hopes that more people push RDNA3 local inference further.
More from Infra
- fal rebuilt MiniMax's open-source H3 video model to run 35x faster, pushing GPUs to 70-80% of ceiling — gorkem · 2026-09-18
- Generating a 2-hour movie for ~$7: heavily optimized MiniMax on Colab — Interesting-Town-433 · 2026-09-18
- Prediction: flagship-quality local models on 16GB machines within 18 months — julianharris · 2026-09-18
- vLLM trains a DSpark speculator for 2.8T-param Kimi K3, hitting ~435 tok/s — AccBalanced · 2026-09-18
- Poll: What LLM gateway do you run at work when every dev holds their own API keys? — almost1it · 2026-09-18
- NVIDIA DGX Station with GB300 demos thousands of tokens per second fully local — TheZachMueller · 2026-09-18