Geek Test: Running a 122B Parameter LLM on an Obsolete Laptop
_TheGreatDreamer_ · reddit · 2026-08-13
A hardcore developer successfully ran a massive 122 billion parameter Qwen model on an obsolete laptop, predating the LLM era, purely for testing purposes.
Although constrained by the hardware, the model took 5 minutes to load, 2 minutes to process the prompt, and 14 minutes to generate a response. However, this extreme test proves the powerful advancements in quantization and inference tools like llama.cpp—showing that running ultra-large models locally is no longer an absolute physical impossibility, even on extremely weak hardware.
More from Infra
- Enthusiasts Discuss Running Massive Qwen3.8-2.4T Models Locally — segmond · 2026-08-13
- Running DeepSeek V4 Flash Locally on 2x DGX Sparks Delivers Prosumer-Grade Performance — andrewchen · 2026-08-13
- Is Local Generative AI Worth It Anymore? Developers Struggle Against Closed Cloud Models — ImaginaryEffective63 · 2026-08-13
- Intel Razor Lake AX Info Surfaces, Targeting AMD's Future Local AI Chips — Terminator857 · 2026-08-13
- Investor Burry Shorts Compute Stocks, Sparking Debate Over AI Compute Shortage — inductionheads · 2026-08-13
- $14.6B AI Compute Bet: Jane Street Needs 20.3% Annual Yield to Break Even — adrianscottcom · 2026-08-13