Achieving 200k context on a single 5090 using Q4 KV Cache and SKILL.state

Ok-Shower7286 · reddit · 2026-08-31

A user running Qwen2.5-32B (Q6) for coding on a single RTX 5090 faced VRAM limits, hitting the ceiling at under 150k context with standard KV Cache.

Solution:

Results:

Original post →

More from coding & agent

coding & agent channel →