85.6 tok/s on a single RTX 5090: local Qwen3 27B with MTP speedup
comperr · reddit · 2026-09-26
A Reddit user reports running Qwen3.8:27B locally on a single RTX 5090, achieving 85.6 tok/s using MTP (multi-token prediction). A useful data point for anyone considering running 27B-class models on consumer hardware.
More from Infra
- MLX-Serve v26.9.6: M5 Ultra optimizations push Qwen3.8 decode to 300+ tok/s locally — TheMoonMidas · 2026-09-26
- AxonDAO launches AxonOS: a GPU-native Linux desktop for science in your browser — melnykowycz · 2026-09-26
- ClusterMAX 3.0 Raises the Bar: Only 19 Neoclouds Get Medallion Ratings, 15 Land in New Bottom Tier — FinanceYF5 · 2026-09-26
- ClusterMAX 3.0 Benchmarks 77 Neoclouds, Tracking 323 Providers Overall — FinanceYF5 · 2026-09-26
- Niantic opens Places Library: 100 real-world 3D scenes for embodied AI training — Scobleizer · 2026-09-26
- umbrelOS 2.0 Launches: Turn a Spare Computer Into a Private Cloud via USB — JosephJacks_ · 2026-09-26