Qwen 3.8 hits 45+ T/ps on M-series chips with detailed optimization guide
Adventurous_Cat_1559 · reddit · 2026-08-21
The author shares optimization experience for running Qwen 3.8 (8-bit quantized) on M-series chips, achieving a stable 45+ T/ps throughput through trial and error, setting a record for this model on the M-series. The post links to a benchmark page with detailed launch arguments, including annotations on what worked and what didn't. The author reports no performance degradation in actual coding tests compared to default settings.
More from Infra
- NVIDIA Details Qwen3.8-2.4T Deployment on GB300, Achieving >4K Tokens/s per GPU — PyTorch · 2026-08-21
- SpaceX launch cadence could enable 12-50 GW of space compute — teortaxesTex · 2026-08-21
- Local AI Coding Hardware Tiers: $1k Gets You the Smartest Model — nickbaumann_ · 2026-08-21
- Cerebras officer Sean Lie files to sell $153M in shares as IPO retail buyers get crushed — firstadopter · 2026-08-21
- Fable launches enterprise safeguards running on your infrastructure for data control — trq212 · 2026-08-21
- IOTA SN9 tests decentralized training of 16B model on mixed GPU clusters — bittingthembits · 2026-08-21