Seeking Real-World Benchmarks: R9 7900XTX Running Qwen2.5-72B

BillyQ · reddit · 2026-08-24

The user plans to buy a PowerColor R9700 (7900XTX) to run Qwen2.5-72B (Q4KXL) via llama.cpp/Vulkan and is seeking real-world token/sec numbers at 64K+ context length under real workloads. Despite AMD's claim of 51.8 tok/s, benchmarks on the NVIDIA 5090 show significant performance degradation for Qwen2.5 as context fills. The user is also inquiring about the stability of MTP speculative decoding to avoid OOMs or garbage output.

Original post →

More from Embodied

Embodied channel →