Quantized open-weight chatbots have run fine on unaccelerated laptops for 2+ years
mattwbaker · x · 2026-10-08
A reminder that you could have been running quantized open-weight chatbots on laptops without any GPU acceleration for over two years — a bit slower, but the author says the experience is pretty great, making local open-source LLM deployment a viable no-cost route.
More from Infra
- llama.cpp merges Metal kernel PR covering all 26 weight formats, up to 4.4x faster MMA on Mac — ggerganov · 2026-10-08
- Theo: viral $132M/year token cost claim is wrong — closer to $3M now, $1.2k soon — dotey · 2026-10-08
- SketchSSM cuts linear-attention state traffic 10x, speeds decode up to 7.3x on B300 — sehoonkim418 · 2026-10-08
- Linear attention takes up to 75% of decode latency at large batch, authors say — sehoonkim418 · 2026-10-08
- Unverified: Baseten Gross Margin at 17%, Cursor Revenue Share Fell from 57% to 28% — menhguin · 2026-10-08
- FT: China races to build AI data centres across energy-rich hinterland — EleanorOlcott · 2026-10-08