DFlash2 Quantization: Qwen3.8 27B on 2x3090 with 262k Context
luedtek · reddit · 2026-08-20
The author released the DFlash2 quantization (W8I) for Qwen3.8-27B. This optimization allows running the model with 262k context at near-full INT8 precision on two 3090 GPUs, achieving speeds of 140 tps. The release is currently in beta, with further implementation optimizations and benchmarks expected.
More from Infra
- GitHub Commits Surge to 1.4B/Month as Cursor Launches Origin Code Hosting — krishnan · 2026-08-20
- Colibri: Pure C engine runs 2.8T parameter MoE models on consumer hardware — tom_doerr · 2026-08-20
- Apple Silicon achieves ANE+GPU dual acceleration, boosting prefill by 50% — bakawolf123 · 2026-08-20
- NVIDIA: Vera Rubin Platform Delivers 10x Tokens/Second per Megawatt — nvidia · 2026-08-20
- Opinion: SpaceXAI emerges as the new No.3 frontier lab with 10GW compute ambitions — thedealdirector · 2026-08-20
- NVIDIA cuOpt becomes fastest open-source solver on Hans Mittelmann benchmarks — NVIDIAAI · 2026-08-20