Ling 3.0 Tiny hits 36 tok/s on 4GB VRAM, matching Qwen 3.5 9B in quality

cosmos_hu · reddit · 2026-08-18

A user tested Ling 3.0 Tiny (8B total, 1.3B active) on an old PC with 4GB VRAM: it generates at 36 token/s, versus 5 token/s for Qwen 3.5 9B on the same machine. The author finds its intelligence very close to Qwen 3.5 9B / Gemma 12, calling it the fastest and smartest model for low-end hardware and hoping for more tiny fast open-source models.

Related event: Ling 3 Tiny Runs at 36 tok/s on 4GB VRAM, Matching Qwen3.5 9B(2 posts)→

Original post →

More from Models

Models channel →