Single AMD R9700 Hits 153 tok/s Running Qwen3.8 27B NVFP4 After Optimization
whodoneit1 · reddit · 2026-09-17
A Reddit user doubled inference performance for a single AMD Radeon R9700 card running Unsloth's Qwen3.8 27B NVFP4: peak decode of 153 tok/s (JSON tasks), 470 tok/s at 8 concurrent requests, and 3,619 tok/s prefill.
Per-category results:
- : 153.1 tok/s (fastest); summarization: 141.7; math: 140; fileedit: 138
- reasoning: 123.9; code: 120.5
- chat: 67.1; prose: 69.2
p50 update latency held steady at 34-42ms across tasks. Results and charts are published in the BetterBench GitHub project — a useful reference for single-GPU AMD setups.
Related event: Single AMD R9700 Hits 153 tok/s on Qwen3.8 27B with NVFP4(2 posts)→
More from Infra
- LinkedIn to present a PyTorch-native GPU retrieval engine powering feed and search — PyTorch · 2026-09-18
- vLLM boosts Kimi K3 serving throughput 2.2-2.8x with scheduler, KDA and MoE kernel optimizations — vllm_project · 2026-09-17
- Jensen Huang: Nvidia expects to sell twice as many chips next year as this year — firstadopter · 2026-09-17
- AWS compares Bedrock RAG vector stores: OpenSearch vs pgvector vs S3 Vectors — AWS ML Blog · 2026-09-17
- Baseten makes pyannote diarization 9.6x faster with 3.2x higher throughput — baseten · 2026-09-17
- huggingface_hub v1.32.0 ships shared download cache, reusing 16.8GB of weights across repos — vanstriendaniel · 2026-09-17