Migrating from LM Studio to vLLM: RTX 3090 performance doubled
Bpthewise · reddit · 2026-08-22
The author shares their experience migrating from LM Studio to vLLM, utilizing configurations from Syv-ai's repository.
Key points:
- Hardware Usage: Running separate endpoints on two RTX 3090s, one for chat and one for sub-agents.
- Performance Gain: Qwen3.8 inference speed significantly increased to 143 tok/s, making the reasoning process much smoother.
- Thermal Improvement: Water-cooled GPUs stay at just 35°C under heavy load, compared to hitting 70°C with LM Studio.
- Conclusion: vLLM feels like fully utilizing the hardware's potential.
More from Infra
- Hiring: Infra engineer role working with xAI Colossus and Oracle Stargate teams — isidentical · 2026-08-22
- Case Study: Therna Bio ships Boltz-2 app in 1 day using Lightning Cloud — LightningAI · 2026-08-22
- $10B Off-Grid DC Possible with Low Battery Prices — aronchick · 2026-08-22
- HBF misunderstood: 3TB/s NAND relies on parallelism, not HBM-like behavior — nickbaumann_ · 2026-08-22
- Marin 535B training starts with full open process and scaling ladder — ysu_nlp · 2026-08-22
- Cloudera launches Anywhere Cloud platform focusing on agents and data sovereignty — DavidLinthicum · 2026-08-22