Tensor-Level Allocation Boosts Qwen 3.5 4B Reasoning by 16.67%
devildip · reddit · 2026-08-22
The author successfully extended tensor-level allocation from the Gemma family to Qwen, achieving significant results on Qwen 3.5 4B.
Using the QLAB allocation strategy on IQ2XS quantization, reasoning performance improved from 46.875 to 54.688—a relative gain of 16.67%—with only a 0.412% increase in model size. The method involves building an imatrix from a category-based corpus, measuring damage, and redistributing precision at the tensor level without post-training or weight updates.
More from Models
- Researcher: GPT 5.6 Sol Ultra Beats Pro for Long-Horizon Hard Problems — arankomatsuzaki · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24