Local benchmark: community 9B model shows 4x lower latency on RTX 5070 vs peers

Storge2 · reddit · 2026-10-02

A Reddit user benchmarked Jev 1.13 (community Jebadiah 9B V2, based on Qwen 3.5 9B) against 4 local LLMs on an RTX 5070 12GB across 20 tasks, reporting roughly 4x lower latency at comparable quality. Small-scale community test, attached screenshot as evidence.

Related event: RTX 5070 Benchmark: 9B Fine-Tuned Model Cuts Latency to a Quarter(2 posts)→

Original post →

More from Models

Models channel →