Indie dev ternary-quantizes Qwen 27B to 7.3GB, keeps image skills for ComfyUI
udmrzn · x · 2026-10-01
- A Japan-based developer (a practicing physician) released Mitsuba, a ternary-quantized, ComfyUI-specialized version of Qwen3.8-27B: 51.5GB shrunk to 7.3GB (14% of original), reportedly the first personal ternary quantization release of its kind in Japan.
- It does one job: look at an image, then write prompts for image/video generation (e.g. with Krea 2 I2T2I). Image recognition drops only 89.8→87.8; fully-compliant generation prompts improve 3/10→6/10; coding collapses 66→4.
- Requires PrismML llama.cpp, Thinking recommended OFF; uncensored version in progress.
- The thread argues purpose-built quantized models may soon let 27B-class VLMs run on 8GB VRAM.
More from Infra
- GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support — ycombinator · 2026-10-01
- Pareto partners with Engy to bring Bittensor SN10 inference optimization to external customers — markjeffrey · 2026-10-01
- Respan launches Span-01 router: 37% cheaper than best single model at matching accuracy — ycombinator · 2026-10-01
- RTX 3090 power can be dialed down to 120W for inference, saving power and heat — QuixiAI · 2026-10-01
- AI as compiler: model writes PTX directly, 1.37x speedup on FlashAttention over Triton — Azaliamirh · 2026-10-01
- Dell ships first Vera Rubin NVL72 rack-scale systems in volume, citing unprecedented NVIDIA partnership — yenkel · 2026-10-01