Postgres memory layer for agents reaches 73.6% on full LongMemEval, not a sampled subset

MycoBrainAI · reddit · 2026-07-21

A Postgres-based memory layer for agents scored 73.6% QA accuracy on the full 500-question LongMemEval oracle set, not a sampled subset.

The builder says the full run exposed problems that the sampled benchmark hid. The biggest improvements came from three choices: handling contradictions explicitly, refusing to store low-confidence extractions, and combining full-text with semantic retrieval instead of using only one. A one-command repro and harness were published in the repo.

Original post →

More from coding & agent

coding & agent channel →