Qwen3.8-27B Benchmarks: RPC inference tests across AMD and Nvidia GPUs

tabletuser_blogspot · reddit · 2026-08-17

The author benchmarked Qwen3.8-27B inference using llama.cpp RPC across various hardware setups. Configurations tested include a solo RX 7900 GRE, a RX 7900 GRE paired with a GTX 1080Ti, and a GTX 1080Ti with dual P102-100 GPUs. Results show that the 16GB RX 7900 GRE struggles with VRAM capacity. With RPC offloading, Q4KM quantization achieved throughput between 71-106 t/s, while Q6K reached 80-116 t/s. Notably, using three GPUs offered no speed advantage over two GPUs, only increasing load times.

Original post →

More from Infra

Infra channel →