Optimizing Qwen3.8 config for single RTX 3090 setups
sagiroth · reddit · 2026-08-16
Reddit users discuss optimal configurations for running the Qwen3.8 model on a single RTX 3090. The conversation covers various inference engines like vLLM and llama.cpp, and tuning settings to maximize performance within VRAM constraints.
More from Infra
- MLX-VLM tops Apple Silicon inference speed with 36.4 tok/s, beating Ollama and llama.cpp — andrejusb · 2026-08-16
- Using If/Else Instead of Agents to Save Costs — newbietofx · 2026-08-16
- AI Compute Speculation: Earth as Computronium for Gaming? — davidmanheim · 2026-08-16
- Does quantizing K/V caches affect model experience? — False-Advantage-4984 · 2026-08-16
- Do we need gateways now that MCPs have gone stateless? — wallphaser231 · 2026-08-16
- Cloudflare Artifacts supports EU and US data localization — dinasaur_404 · 2026-08-16