Qwen 3.5 122B Local Quantization Tested
RedParaglider · reddit · 2026-07-15
Trying this for the first time, the author ran Qwen 3.5 122B Heretic ROCmFP4 iMatrix on Strix Halo based on their needs, sharing the benchmark results:
- Total parameters: 122B, active parameters: 10B
- VRAM usage: 60.70 GiB
- Speed: 28.45 tok/s
- BF16 KLD: 0.100716
The author notes that if there's demand for a non-Heretic version, they will release it when compute allows.
Related event: Local Benchmark: Qwen3.5 122B ROCmFP4 Quantization(2 posts)→
More from Models
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22