Benchmarking Qwen3.8 BF16 across RTX 5090 and Mac via RPC

stargate425 · reddit · 2026-08-31

Ran Qwen3.8-Flash-Next in full BF16 via llama.cpp RPC across RTX PRO 6000, RTX 5090, and Mac Pro. Achieved 6.43 t/s decode speed. Investigating RPC bottlenecks as 39GB spills to RAM and performance lags behind single-node setups.

Original post →

More from Infra

Infra channel →