Developer Achieves 114 TPS on 30B Model Using Quantization and Slicing
dejavucoder · x · 2026-08-10
A developer shared a successful use case of running a large language model locally in a reply. By utilizing the Muse Glimmer 30B model combined with specific quantization (tuned quants) and dflash techniques, they achieved a high inference speed of 114 tokens/s.
The developer noted that this result is promising, expressing plans to further slice out the model's agentic capabilities to make it run smoothly on consumer hardware with less than 16GB of VRAM.
Related event: Developer Achieves 114 TPS Running 30B Model Locally(2 posts)→
More from Infra
- Muse Glimmer Hits 230 tok/s on a Single RTX 5090 via SGLang — BanghuaZ · 2026-08-11
- SGLang v0.5.17 Released: Adds Support for Kimi K3 and MiniMax-H3 Video Generation — BanghuaZ · 2026-08-11
- Karpathy: Hybrid Setup of Cloud Executive Intelligence and Local Models is Very Appealing — karpathy · 2026-08-11
- Local Video Gen with MiniMaxH3: Workflow and Hardware Upgrade Notes — Last-Pie8057 · 2026-08-10
- Running Large Models on 8GB VRAM? Understanding Shared Memory in Local Deployment — Leary_2844 · 2026-08-10
- SemiAnalysis Deep Dive: Can TileRT Software Make NVIDIA GPUs Compete with Cerebras and Groq? — dylan522p · 2026-08-10