Custom Gem16 engine runs Gemma4 26B at 182 tok/s on a 16GB laptop GPU

Danmoreng · reddit · 2026-09-28

Developer Danmoreng built Gem16, a small custom inference engine for Gemma4 on Blackwell 16GB GPUs — mostly vibe-coded with Codex but tuned over many weekends to beat vLLM, which couldn't handle MTP in 16GB VRAM. Runs on Linux and Windows, has a native GUI, and targets single-user serving (the 12B model serves two sessions).

Performance:

Known issue: the 12B's native audio understanding stops recognizing audio tokens after 8k context — apparently a model-side bug others have reported. Open-sourced on GitHub, feedback welcome.

Original post →

More from Infra

Infra channel →