Benchmarking Qwen3.8 BF16 across RTX 5090 and Mac via RPC
stargate425 · reddit · 2026-08-31
Ran Qwen3.8-Flash-Next in full BF16 via llama.cpp RPC across RTX PRO 6000, RTX 5090, and Mac Pro. Achieved 6.43 t/s decode speed. Investigating RPC bottlenecks as 39GB spills to RAM and performance lags behind single-node setups.
More from Infra
- OpenAI reportedly buying tens of thousands of Macs for RL training — SumitGup · 2026-08-31
- Data center reporting should learn from telegraph cable narratives — jwt0625 · 2026-08-31
- New project gpu_api introduces minimal API design, optimizing graphics rendering experience — Michael_Moroz_ · 2026-08-31
- Hugging Face used open weights to defend; frontier APIs failed — AlexTensor · 2026-08-31
- Deep Dive into the Machine God's Blood Supply: AI Compute Economy Explained — ZeroStateReflex · 2026-08-31
- Paper runs distributed LLM inference over 10 km of multi-core fiber, no simulation — jwt0625 · 2026-08-31