Qwen3.8-27B MTP Quants on M5 Max: oQ4e Best Balance
DerTomsn · reddit · 2026-08-31
Author benchmarked Qwen3.8-27B MTP from oQ2e to oQ8e on Apple M5 Max. Results show: oQ2e is unusable; oQ3e offers good quality; oQ4e is the sweet spot balancing speed (36.6 tok/s) and quality (86.8); oQ8e delivers peak quality (87) but with slowest speed and 2x memory footprint.
More from Models
- Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization — alejandroll10 · 2026-09-01
- Open Source Models Shift to Revenue Sharing and Licensing — zephyr_z9 · 2026-09-01
- Has anyone tuned a model to operate exclusively in E-prime yet? — cephaloform · 2026-09-01
- Heavy users report Claude quality dropping over the past week: eager to execute, no more clarifying questions — Numerous_Leopard_522 · 2026-09-01
- User seeks best LLM for CLI coding on single 3080 Ti — -samae1- · 2026-09-01
- Lan Hackathon Review: Qwen Excels in Physics/Engineering, K3 in General Intelligence — 葬AI · 2026-09-01