DFlash2 Quantization: Qwen3.8 27B on 2x3090 with 262k Context

luedtek · reddit · 2026-08-20

The author released the DFlash2 quantization (W8I) for Qwen3.8-27B. This optimization allows running the model with 262k context at near-full INT8 precision on two 3090 GPUs, achieving speeds of 140 tps. The release is currently in beta, with further implementation optimizations and benchmarks expected.

Original post →

More from Infra

Infra channel →