Ling Tiny 3.0 on a 2017 laptop: 8B model hits 10 tok/s, hints at edge AI future

netherreddit · reddit · 2026-09-26

A Reddit user ran the 8B-parameter Ling 3.0 Tiny with llama.cpp on a 2017 laptop (7th-gen i5, 8GB RAM, no GPU) at 10 tokens/s. In 20 minutes the model completed a multi-turn coding task: writing, running, and iterating on a script that scans a local network for llama.cpp servers. The author argues that 1B active parameters now make useful on-device intelligence possible on nearly any post-2015 hardware, potentially kicking off a new era of edge intelligence with zero new compute.

Original post →

More from Infra

Infra channel →