RTX 3090 gets 35 t/s on Qwen 3.8 27B
cviperr33 · reddit · 2026-08-15
A user reports running Qwen 3.8 27B (IQ4 NL quant) on an RTX 3090 via llama.cpp at 34-36 tokens/s. This is slower than the 50-60 t/s they recalled getting on v3.6, without speculative decoding enabled. Despite the speed dip, they feel the model quality is truly next-gen.
More from Infra
- SpaceX partners with Nvidia on orbital data centers, first satellite launching next year — JOBhakdi · 2026-08-15
- Qdrant + Minima Boost Agentic RAG 2.92x on Single RTX PRO 6000 — qdrant_engine · 2026-08-15
- NVIDIA open-sources NeMo Switchyard for dynamic model routing in agent workflows — NVIDIAAI · 2026-08-15
- Vercel ranked as the world's fastest AI Gateway infrastructure — cramforce · 2026-08-15
- Qwen3.8-2.4T-A95B deployment guide: NVFP4 needs 8×B300, TP must divide 64 — Necessary_Gazelle211 · 2026-08-15
- mcpp: Auto-generate MCP servers from C++ code via reflection — karurochari · 2026-08-15