Custom inference engine pushes Qwen3.8-Flash-Next to 65 tok/s on 12GB RTX 5070
KnownAd4832 · reddit · 2026-09-25
A Redditor built a custom inference engine (Strata, open source, one-click install, CUDA-optimized) after getting only 15 tok/s with llama.cpp on a 12GB RTX 5070. The same IQ3XXS quant of Qwen3.8-Flash-Next now runs at 65 tok/s output and 430 tok/s prompt processing. On 128K context with 64GB DDR5 + Ryzen 5 7600: Q20 at 65.1 tok/s out / 543 pp, IQ2XS at 52.0 / 472, IQ3XXS at 44.8 / 414. Minimum RAM+VRAM: 37.6GB (Q20) to 47GB (IQ3XXS), plus 0.91GB for the vision encoder. Model: ISTA-DASLab GGUF quants.
More from Infra
- Investor Predicts EDA/CAD Will Collapse Into One Flow Within 3-5 Years — ai · 2026-09-25
- AI Data Center Debt Starting to Roll Over, Rising Rates Accelerating the Problem — AIFlow_ML · 2026-09-25
- Musk details xAI compute: Colossus 2 to hit 880k GB300s by year-end — elonmusk · 2026-09-25
- Qwen-Image-2.1 gets GGUF quantization, could run text-to-image on a Snapdragon 865 phone — ResidentAping · 2026-09-25
- Deep Inference-Query Engine Integration: Custom Scheduler and Workload-Aware KV Cache for Prefill-Only AI Filters — charles_irl · 2026-09-25
- Finance worker seeks local AI setups to cut soaring Codex/ChatGPT costs — Startup__Sam · 2026-09-25