200 tok/s on 8GB VRAM: dev benchmarks 6 small models for local AI

TheMoonMidas · x · 2026-09-06

Developer @nettermina shows a max-optimized local setup: no $15,000 GPU needed — an old 8GB VRAM card (tested on a forgotten 3070 Ti laptop) runs Windows with 200 tok/s, 128k context, and text/image/audio support.

The team pulled every instruct model since May that fits 8-12GB and ran six through the same 20 tasks (JSON extraction, code, tables, arithmetic, deep-prompt fact recall, tool calls, photo reading): Gemma 4 E4B, Gemma 4 12B, Qwen3.5-9B, Qwen3.5-4B, Ornith-1.5-9B, and a Qwen3.8-9B distill.

Result: Gemma 4 E4B isn't the top scorer — Qwen3.5-9B and Ornith beat it on long context and fine print in photos — but E4B won the pick for what a newcomer on an old card actually needs. A reproducible recipe is included for anyone with a 5-year-old laptop.

Original post →

More from Infra

Infra channel →