RTX 3090 Runs 27B Model with 200K Context via DFlash2
MaziyarPanahi · x · 2026-09-01
MiaAI Lab released an EXL3 quantization deployment kit for Qwen3.8-27B, enabling 200K context on 24GB VRAM (e.g., RTX 3090) using DFlash2 speculative decoding. The repo includes launchers, configs, and an OpenAI-compatible server with NVFP4 KV cache for efficiency.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- 2 engineers + AI designed a working LLM chip in 2 weeks, no human in the loop — 新智元 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01