Developer Achieves 114 TPS on 30B Model Using Quantization and Slicing

dejavucoder · x · 2026-08-10

A developer shared a successful use case of running a large language model locally in a reply. By utilizing the Muse Glimmer 30B model combined with specific quantization (tuned quants) and dflash techniques, they achieved a high inference speed of 114 tokens/s.

The developer noted that this result is promising, expressing plans to further slice out the model's agentic capabilities to make it run smoothly on consumer hardware with less than 16GB of VRAM.

Related event: Developer Achieves 114 TPS Running 30B Model Locally(2 posts)→

Original post →

More from Infra

Infra channel →