RTX 3090 Runs 27B Model with 200K Context via DFlash2

MaziyarPanahi · x · 2026-09-01

MiaAI Lab released an EXL3 quantization deployment kit for Qwen3.8-27B, enabling 200K context on 24GB VRAM (e.g., RTX 3090) using DFlash2 speculative decoding. The repo includes launchers, configs, and an OpenAI-compatible server with NVFP4 KV cache for efficiency.

Original post →

More from Infra

Infra channel →