ISTA-DASLab splits prefill/decode quantization, squeezes 27B-class model into 13.7GB GGUF

victormustar · x · 2026-09-21

ISTA-DASLab released an experimental HF repo, Qwen3.8-27B-disaggregated-NVFP4-prefill, premised on the idea that prefill and decode have different bottlenecks and needn't share one quantization format:

It remains experimental: no model card, custom odp architecture, and not a drop-in GGUF for existing runtimes.

Original post →

More from Infra

Infra channel →