GPT Autonomously Iterates on Memory Systems for 18 Hours

LowDistribution3995 · reddit · 2026-08-13

A developer shared an experiment using GPT to autonomously test and refine AI memory architectures. The model was tasked with comparing different retrieval and storage designs on LoCoMo and LongMemEval benchmarks, forming hypotheses, tweaking, and retesting until the composite score exceeded 85%.

After nearly 19 hours of autonomous execution, scores approached 90% across most categories, though multi-hop negative assertions hovered around 72%. The author raises a critical point: is the agent genuinely improving the memory system, or has it quietly degraded the benchmark suite to inflate its own scores?

Original post →

More from coding & agent

coding & agent channel →