Qwen3.8-27B Ridge: Smarter quantization released
alexcovo_eth · x · 2026-08-19
Qwen3.8-27B Ridge is a new quantized version maintaining 27B intelligence while fitting in 12GB VRAM on a single consumer GPU. Key features include:
- Architecture-aware quantization: GDN state protected at Q80 to avoid precision loss of generic IQ2.
- Built-in MTP speculative decoding: Faster tokens with zero quality loss.
- Multimodal support: Works with mmproj for images (paintings, signs, etc.).
- Benchmarks: Tested with Mark Twain style writing, tree animation, and 80-language test, all running locally on one card.
More from Infra
- Hugging Face docs update: Shipping Triton autotune configurations — RisingSayak · 2026-08-19
- Matryoshka LM Suites: Train model suites in a single run — yoavartzi · 2026-08-19
- Liquid releases QAD quantized models, recovering 97% accuracy at 4-bit — helloiamleonie · 2026-08-19
- LoRA Training Test: Gradient Accumulation Is Not a Time-Zero Game — traceml-ai · 2026-08-19
- CXL standard nears deployment, benefiting Marvell and Google — BenBajarin · 2026-08-19
- LLM Midtraining Guide: Data Mixing and Overfitting Prevention — cwolferesearch · 2026-08-19