Memory layer cuts agent context tokens 23-62x and beats full history on 90-day recall

No_Advertising2536 · reddit · 2026-09-16

A memory API builder published a rare measured A/B on Reddit: across three synthetic dialogue corpora (companion, support, coding) with planted facts and ground truth, comparing a memory layer (per-day fact extraction into a 600-token budget) vs sending full history, with gpt-4o-mini as the answer model.

Key numbers

Three surprises

The bench, corpus generator (fixed seed), and cost report are all open-sourced.

Original post →

More from coding & agent

coding & agent channel →