8B model trained on distributed gaming GPUs for $6,500 runs on phone CPU at ~60 tok/s
markjeffrey · x · 2026-10-01
Jon Durbin of Chutes.ai presented at Exploit Summit: an 8B model trained across distributed gaming GPUs for roughly $6,500 in GPU rental, running entirely on a phone's CPU at nearly 60 tokens per second. He argued open-source development and decentralized training could give people control over AI from creation to daily use — building toward a "Linux of AI."
More from Infra
- GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support — ycombinator · 2026-10-01
- Pareto partners with Engy to bring Bittensor SN10 inference optimization to external customers — markjeffrey · 2026-10-01
- Respan launches Span-01 router: 37% cheaper than best single model at matching accuracy — ycombinator · 2026-10-01
- RTX 3090 power can be dialed down to 120W for inference, saving power and heat — QuixiAI · 2026-10-01
- AI as compiler: model writes PTX directly, 1.37x speedup on FlashAttention over Triton — Azaliamirh · 2026-10-01
- Dell ships first Vera Rubin NVL72 rack-scale systems in volume, citing unprecedented NVIDIA partnership — yenkel · 2026-10-01