Unlock 250k Context on RTX 4090: Optimizing Qwen 3.8 with Custom Drafters

alexcovo_eth · x · 2026-08-23

The author achieved a 250,000 token context window for Qwen 3.8-27B on a single RTX 4090 (24GB) running at 75 tokens/s through specific optimizations.

Key Optimization Techniques:

Results:

Combining the VRAM savings from the Q2K drafter with the --parallel 1 flag caused the context ceiling to "absolutely explode," enabling the massive 250k context length on consumer hardware.

Original post →

More from Infra

Infra channel →