Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU

MoonsvnLyn · reddit · 2026-10-04

On a 5070 Ti (16GB) + 14700KF + 96GB RAM, the author ran Qwen3.8-Flash-Next IQ3S with 262K context through Strata (KV cache in RAM streaming a 32K window into VRAM), lifting decode from 17.2 tok/s to 43 cold and 53.5 with prefix cache — with full repro steps:

Caveat: re-running setup rewrites the config and drops profile flags. Full numbers live in the author's tuning doc on the Strata repo, plus an upstream issue filed for the unmeasured IQ3S@262K data point.

Original post →

More from coding & agent

coding & agent channel →