Qwen3.8-27B at 256K Hits 50 tok/s on a 24GB RTX 4000 SFF
pich · hn · 2026-08-17
- Setup: 24GB RTX 4000 SFF (432 GB/s bandwidth).
- Config: Qwen3.8-27B model with 256K context window enabled.
- Result: Achieved 50 tokens/second throughput using MTP (Multi-Token Prediction).
- Details: The post analyzes how VRAM bandwidth impacts inference speed and demonstrates practical gains from MTP.
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24