Claim: run a 125B MoE at 22 tok/s on $320 of used GPUs with llama.cpp

cephaloform · x · 2026-09-16

A user claims their buun-llama-cpp setup runs a 125B MoE ("Qwen3.8-Flash-Next") at Q4 quantization with 22.3 tok/s sustained on just $320 of GPUs ($80 P100s), calling it the pareto frontier of speed/intelligence/cost. The model name and numbers are unverified, but the cheap used-GPU approach is noteworthy if real.

Original post →

More from Infra

Infra channel →