Claim: run a 125B MoE at 22 tok/s on $320 of used GPUs with llama.cpp
cephaloform · x · 2026-09-16
A user claims their buun-llama-cpp setup runs a 125B MoE ("Qwen3.8-Flash-Next") at Q4 quantization with 22.3 tok/s sustained on just $320 of GPUs ($80 P100s), calling it the pareto frontier of speed/intelligence/cost. The model name and numbers are unverified, but the cheap used-GPU approach is noteworthy if real.
More from Infra
- Dev slams agent sandbox pricing as 20x+ the cost of a $6/month always-on VPS — Aryvyo · 2026-09-16
- Signal65 benchmarks NVIDIA Vera CPU at 1.64x per core vs x86 on agentic workloads — ryanshrout · 2026-09-16
- McKinsey: Sovereign AI TAM to hit $500-600B by 2030, 30-40% of AI demand — Beth_Kindig · 2026-09-16
- Intel CEO Lip Bu Tan tells AI Infra Summit: embrace AI, don't fear it — karlfreund · 2026-09-16
- Brad Gerstner: AI buildout hinges on revenue keeping its steep growth curve — markjeffrey · 2026-09-16
- Reported: NVIDIA hardware delivers 40% more throughput at the same power — karlfreund · 2026-09-16