Qwen3.8-27B Benchmarks: RPC inference tests across AMD and Nvidia GPUs
tabletuser_blogspot · reddit · 2026-08-17
The author benchmarked Qwen3.8-27B inference using llama.cpp RPC across various hardware setups. Configurations tested include a solo RX 7900 GRE, a RX 7900 GRE paired with a GTX 1080Ti, and a GTX 1080Ti with dual P102-100 GPUs. Results show that the 16GB RX 7900 GRE struggles with VRAM capacity. With RPC offloading, Q4KM quantization achieved throughput between 71-106 t/s, while Q6K reached 80-116 t/s. Notably, using three GPUs offered no speed advantage over two GPUs, only increasing load times.
More from Infra
- PufferLib prerelease supports >20M steps/second reinforcement learning — jsuarez · 2026-08-17
- Data center boom drives up emissions as costs outweigh benefits — robleclerc · 2026-08-17
- DeepSeek v4 PRO Q2 hits 45 t/s with VRAM/RAM split on DGX Station — antirez · 2026-08-17
- Model routing should be optimized at the harness layer, not the gateway — agihouse_org · 2026-08-17
- AI Spending Gap 625x: Top 1% Spend $7,500/Employee/Month vs Median $12 — rohanpaul_ai · 2026-08-17
- AWS to deploy Nvidia GB300 as primary GPU in 2026, expand Trainium shipments: TrendForce — Beth_Kindig · 2026-08-17