Strata on a 4090 48GB: conversation parking buys 35x, n-gram table in RAM just 1%

Shot-Ad-4147 · reddit · 2026-10-04

A detailed single-machine benchmark of Strata 0.1.38 on i9-13900K + RTX 4090 48GB + 128GB DDR5, using a real 91,836-token prompt.

Four-arm test of n-gram table placement

Methodology lesson: the author's earlier +4.7% claim came from a wrong baseline — session-to-session spread on identical config was 11.5% (149.1 → 166.3 tok/s). A/B runs must happen back-to-back in the same session.

Biggest win — conversation parking (35x): with --conversation-cache-mib 8192 and 4 slots, returning to a 91,836-token conversation drops from 19.5s to 537 ms for just 3.1 GB of RAM.

Other wins: --prefill auto:32768 gives +9.7% prefill / +13.6% decode at 78.7K context; --calibrate found another +3.2% by changing CPU workers 23 → 12. Baseline throughput: 128-151 tok/s decode, 4,249-5,154 tok/s prefill.

Original post →

More from Infra

Infra channel →