200 tok/s on 8GB VRAM: dev benchmarks 6 small models for local AI
TheMoonMidas · x · 2026-09-06
Developer @nettermina shows a max-optimized local setup: no $15,000 GPU needed — an old 8GB VRAM card (tested on a forgotten 3070 Ti laptop) runs Windows with 200 tok/s, 128k context, and text/image/audio support.
The team pulled every instruct model since May that fits 8-12GB and ran six through the same 20 tasks (JSON extraction, code, tables, arithmetic, deep-prompt fact recall, tool calls, photo reading): Gemma 4 E4B, Gemma 4 12B, Qwen3.5-9B, Qwen3.5-4B, Ornith-1.5-9B, and a Qwen3.8-9B distill.
Result: Gemma 4 E4B isn't the top scorer — Qwen3.5-9B and Ornith beat it on long context and fine print in photos — but E4B won the pick for what a newcomer on an old card actually needs. A reproducible recipe is included for anyone with a 5-year-old laptop.
More from Infra
- T-Glass shortage worsens: Kinsus losing 10-15% of monthly ABF revenue, 25% capacity expansion planned for 2027 — zephyr_z9 · 2026-09-06
- Bump-less 3D stacking goes practical: Intel Diamond Rapids first, AMD Zen rumored next — bookwormengr · 2026-09-06
- Self-hosting AI: what rigs do local LLM runners actually use? — Fun_Kangaroo512 · 2026-09-06
- Reddit debate: is compute the real hurdle to automating all cognitive labour? — Vivid-Flamingo-644 · 2026-09-06
- Reddit debate: is compute the real hurdle to automating all cognitive labour? — Vivid-Flamingo-644 · 2026-09-06
- Block KV cache streaming PR bounds VRAM at long context in llama.cpp fork — giveen · 2026-09-06