PAST-Bench: A New Benchmark for Evaluating Agent Memory and Persistence

alifcoder · x · 2026-08-05

PAST-Bench is a new benchmark designed to evaluate whether personal agents actually improve from past experience. It features 26 task-family scenarios and 204 episodes, testing 4 key capabilities: Memory, Procedural Reuse, Information Gathering, and Update.

Evaluating across 7 base models and 4 agent frameworks, the study reveals that while retained experience can improve future behavior, the gains are highly capability-, model-, and framework-dependent. This indicates that builders should test memory designs on their own stacks rather than blindly copying them.

Related event: Princeton Introduces PAST-Bench to Evaluate Personal Agent Evolution(3 posts)→

Original post →

More from coding & agent

coding & agent channel →