RTX 3060 Test: ByteShape Ultra-Low-Bit Quant Far Behind GSQ Despite Paper Similarity
zyxciss · reddit · 2026-09-23
On an RTX 3060 12GB + 16GB single-channel DDR4 (CachyOS, llama.cpp), the author compared two ultra-low-bit quants of Qwen 3.8 27B:
- ISTA-DASLab GSQ-RCO-IQ3-XXS (10.4GB, 2.5 BPW, MTP enabled): 29 tok/s at full context, 34-40 tok/s lower;
- ByteShape IQ3-XXS: 500MB smaller, advertised as extremely close to BF16 by KL-divergence.
In practice the gap was large: GSQ generated a 3D voxel diorama in one shot under 55k tokens, while ByteShape took 3 attempts and 98k+ tokens without finishing; web dev results also fell short of the advertised benchmarks. The takeaway: paper similarity metrics for low-bit quants don't guarantee real task performance. The author also asks for model suggestions for RTX 3060.
More from Infra
- DeepSeek Elastic Compute Paper Reveals 160-Node Units Serving 3M Sandboxes Daily — zephyr_z9 · 2026-09-23
- SAFi: an open-source local AI application layer you can pair with llama.cpp for business deployments — forevergeeks · 2026-09-23
- Claude Opus 5.5 lands on Amazon Bedrock with lower pricing and built-in safety classifiers — AWS ML Blog · 2026-09-23
- Ben Bajarin launches expert interview series on powering AI datacenters with gas — BenBajarin · 2026-09-23
- umbrelOS 2.0 launches, turning a small home computer into your personal cloud — JosephJacks_ · 2026-09-23
- fal's H3 Max generates 5 seconds of frontier-quality video in 3 seconds via full-stack optimization — gorkem · 2026-09-23