156-run benchmark: RTX 3070 power limits vs local LLM efficiency, Gemma4 wins
bulletrhli · reddit · 2026-09-27
A data engineer ran 156 benchmarks (4 models × 6 power limits from 100-220W × 3 runs) on a Lenovo M920Q with an i5-8500T and an RTX 3070 8GB over OcuLink, using Proxmox LXC + Ollama + OpenWebUI, logging tokens/s, duration, and a tokens-per-watt metric.
Key findings:
- gemma4:e4b is the best daily driver: no overthinking, controllable tone, and room for 16k context
- qwen3.5 produces good output but is a heavy thinker, briefly hitting 203W
- qwen3-vl over-thinks and balloons context, often knowing the answer early but looping; the author may go back to traditional OCR instead of reasoning vision models
- deepseek-coder-v2 stays fast at a 40/60 CPU/GPU split thanks to its MoE architecture
Below a threshold, the power limit acts as the ceiling—useful guidance for anyone tuning local deployments on tight hardware.
More from Infra
- OpenAI files trademark for "Scotch Bonnet", covering AI processors, HPC and data center hardware — imjustnewatai · 2026-09-27
- OpenAI files trademark for "Cayenne" covering AI chips and data center hardware — imjustnewatai · 2026-09-27
- OpenAI Files 'Serrano' Trademark Covering AI Chips and Data Center Hardware — imjustnewatai · 2026-09-27
- NVIDIA B300 index goes live at $6.95/GPU-hr, rents up 43.9% since April — AccBalanced · 2026-09-27
- ENGRAM and YOCO memory architectures could shift data center DRAM/HBM economics — AccBalanced · 2026-09-27
- Space Compute to Cost 15-30% More Than Terrestrial by 2028-30, but Near-Zero Opex — JOBhakdi · 2026-09-27