Custom Gem16 engine runs Gemma4 26B at 182 tok/s on a 16GB laptop GPU
Danmoreng · reddit · 2026-09-28
Developer Danmoreng built Gem16, a small custom inference engine for Gemma4 on Blackwell 16GB GPUs — mostly vibe-coded with Codex but tuned over many weekends to beat vLLM, which couldn't handle MTP in 16GB VRAM. Runs on Linux and Windows, has a native GUI, and targets single-user serving (the 12B model serves two sessions).
Performance:
- Gemma4 12B with audio+vision: 5,800 tok/s prefill, 87 tok/s decode
- Gemma4 26B with vision, custom EXL3-like quantization: 5,660 tok/s prefill, 182 tok/s decode, fits 220k context
Known issue: the 12B's native audio understanding stops recognizing audio tokens after 8k context — apparently a model-side bug others have reported. Open-sourced on GitHub, feedback welcome.
More from Infra
- Caching 20k-Token Pi Prompts Across Sessions With llama.cpp Slots — ea_man · 2026-09-28
- Spectral deflation framework improves Muon: consistent validation loss gains in GPT-2 pretraining — hankyang94 · 2026-09-28
- First Audited Look at Inference-Economics: MiniMax Hit 24.6% Margin, Peer Lost 75% of OpenRouter Volume — AccBalanced · 2026-09-28
- Yunnan Germanium report: indium for InP is tight in China, export controls aren't the bottleneck — pstAsiatech · 2026-09-28
- Apple's free on-device fm paired with decision model Jev beats big-model routing in tests — jasonkneen · 2026-09-28
- Estimate: crudely describing human biology needs 1000x more data than humanity stores — IgorCarron · 2026-09-28