Benchmark: Qwen 3.8 27B hits 50-60 t/s on dual RTX 5060 TI cards
chocofoxy · reddit · 2026-08-18
A user reported that running Qwen 3.8 27B in Q4 quantization on dual RTX 5060 TI 16GB cards achieves 50-60 tokens/sec (MTP). This is surprisingly faster than the 30 t/s observed with Qwen 3.6, likely due to optimization differences or the switch from Windows to Linux.
More from Models
- Anthropic's next Mythos model reportedly shows only modest improvement — haider1 · 2026-08-18
- Google One vs. Code Assist: small business owner puzzles over the 1,500 daily request cap — innovaldragon · 2026-08-18
- Too Smart to Work With: Users Downgrade to Dumber Models for Readable Output — AashaySachdeva · 2026-08-18
- Muse Spark 1.2 release claims cost-intelligence Pareto frontier — zainhas · 2026-08-18
- LiquidAI releases LFM2.5-VL-3B-WebGPU, a multimodal model optimized for browser deployment — LiquidAI · 2026-08-18
- McByte Sets New SOTA on SportsMOT Benchmark — NielsRogge · 2026-08-18