30B Model Hits 114 tps with Tuned Quants, Targeting <16G VRAM

tokenbender · x · 2026-08-10

Developer @tokenbender shared recent benchmarks in local LLM inference optimization. By utilizing tuned quants and dflash, the Muse Glimmer 30B model achieved an impressive inference speed of 114 tokens/s.

This progress makes him confident that he can eventually slice out the model's agentic capabilities and run it smoothly on consumer hardware with less than 16GB of VRAM.

Original post →

More from Infra

Infra channel →