Ling 3.0 Tiny hits 36 tok/s on 4GB VRAM, matching Qwen 3.5 9B in quality
cosmos_hu · reddit · 2026-08-18
A user tested Ling 3.0 Tiny (8B total, 1.3B active) on an old PC with 4GB VRAM: it generates at 36 token/s, versus 5 token/s for Qwen 3.5 9B on the same machine. The author finds its intelligence very close to Qwen 3.5 9B / Gemma 12, calling it the fastest and smartest model for low-end hardware and hoping for more tiny fast open-source models.
Related event: Ling 3 Tiny Runs at 36 tok/s on 4GB VRAM, Matching Qwen3.5 9B(2 posts)→
More from Models
- User frustrated: Claude spends 2 minutes reasoning on trivial tasks — nikvassev · 2026-08-19
- Qwen3.8 vs 27B: Slower tokens but faster, better results — surreal_tournament · 2026-08-19
- The Trick Behind Agentic Models: Always Falling Forward — But Restraint Gets Harder — generativist · 2026-08-19
- DeepSeek V4 'J-Space' Framework Exposed as Fake: Community Benchmarks Contradict Claims — bookwormengr · 2026-08-19
- Stats: DeepSeek Dominates Token Usage on AIWayfinder Platform — templecrash · 2026-08-18
- DeepSeek v4 local deployment hits 3000 t/s prefill on DGX Station — antirez · 2026-08-18