LM Studio Slower Than Raw llama.cpp by ~40% on Qwen
MkGod · reddit · 2026-08-24
User benchmarked Qwen3.8-27B on dual RTX 5060 Ti, finding a significant performance gap between LM Studio and raw llama-server.exe.
Performance Delta:
- LM Studio: 30–40 tok/s.
- Raw llama-server.exe: 50–55 tok/s (40% faster).
- Settings were identical (TP, MTP speculative decoding), ruling out VRAM spillover.
Analysis:
- Suspects GUI wrapper overhead, context shift, or KV cache fragmentation in LM Studio.
- Despite a warning about CPU sampling fallback in SPLITMODETENSOR, raw llama.cpp remains significantly faster.
Reasoning Effort:
- lmstudio-community GGUF shows full Reasoning Effort dropdown.
- Unsloth GGUF shows only a basic On/Off toggle.
- Question: Is this controlled by GGUF metadata/Jinja templates or hardcoded LM Studio logic?
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24