Quantized Qwen3.8 draft model cuts VRAM 2.6x, boosts context from 90K to 130K on one RTX 5090

QuixiAI · x · 2026-08-30

Original post →

More from Infra

Infra channel →