Local LLM benchmark: fine-tuned 9B model cuts latency 4x on RTX 5070 at same quality

Storge2 · reddit · 2026-10-02

A Reddit user benchmarked a local fine-tune of Qwen 3.5 9B (Jebadiah 9B V2) against 4 other local LLMs on an RTX 5070 12GB across 20 tasks.

The fine-tuned 9B delivered 4x lower latency at equivalent quality, which the author found impressive.

Related event: RTX 5070 Benchmark: 9B Fine-Tuned Model Cuts Latency to a Quarter(2 posts)→

Original post →

More from Models

Models channel →