Tensor-Level Allocation Boosts Qwen 3.5 4B Reasoning by 16.67%

devildip · reddit · 2026-08-22

The author successfully extended tensor-level allocation from the Gemma family to Qwen, achieving significant results on Qwen 3.5 4B.

Using the QLAB allocation strategy on IQ2XS quantization, reasoning performance improved from 46.875 to 54.688—a relative gain of 16.67%—with only a 0.412% increase in model size. The method involves building an imatrix from a category-based corpus, measuring damage, and redistributing precision at the tensor level without post-training or weight updates.

Original post →

More from Models

Models channel →