Ling Tiny 3.0 on a 2017 laptop: 8B model hits 10 tok/s, hints at edge AI future
netherreddit · reddit · 2026-09-26
A Reddit user ran the 8B-parameter Ling 3.0 Tiny with llama.cpp on a 2017 laptop (7th-gen i5, 8GB RAM, no GPU) at 10 tokens/s. In 20 minutes the model completed a multi-turn coding task: writing, running, and iterating on a script that scans a local network for llama.cpp servers. The author argues that 1B active parameters now make useful on-device intelligence possible on nearly any post-2015 hardware, potentially kicking off a new era of edge intelligence with zero new compute.
More from Infra
- steipete pushes back on agent cost concerns: capable models now $0.1 per million tokens — steipete · 2026-09-26
- e/acc leader shows off e/acc-wrapped home inference cluster, dubs it the new status symbol — beffjezos · 2026-09-26
- US AI buildout projected at $10.3 trillion over 6 years, bigger than rail, highways, electrification and telecom combined — 141_1337 · 2026-09-26
- Oh My Pi adds custom model backends: vLLM, llama.cpp, SGLang and more — bolts98 · 2026-09-26
- OpenRouter launches typesafe/jev-router, a cache-aware router that picks models per request — alexcovo_eth · 2026-09-26
- Akamai beats neo-clouds on profitability, Anthropic deal and buybacks, argues investor — pdamodaran · 2026-09-26