Ling 3.0 Tiny vs Gemma 26B-A4B: 3x Smaller VRAM, But Accuracy Halved

autonoma_2042 · reddit · 2026-09-21

In an audiobook pipeline (attributing dialogue speakers for per-character TTS), a developer ran an A/B test of Ling 3.0 Tiny (7.9B total / 1.3B active, 4.8GB at Q4, fits fully in 8GB VRAM) against the incumbent Gemma 4 26B-A4B (QAT Q4, 14GB, experts in system RAM):

Results: on 485 hand-corrected lines, Gemma scored 94.6% overall (worst chapter 90.6%) vs Ling's 56.9% (worst 43.8%), with the gap widening on longer chapters; Gemma was also faster wall-clock.

Takeaway: Ling 3.0 Tiny's memory and throughput advantages don't translate to quality on strict JSON extraction — the 26B MoE wins decisively.

Original post →

More from Models

Models channel →