PAST-Bench: Evaluating Recursive Self-Improvement in Personal AI Agents
rohanpaul_ai · x · 2026-08-05
Researchers introduced PAST-Bench, a benchmark designed to test whether personal AI agents can turn accumulated experience into better future behavior.
- Scale: Spans 26 scenarios and 204 episodes across 4 core capabilities: memory, procedural reuse, information gathering, and update.
- Methodology: Evaluates agents under matched conditions with experience retention turned on/off, tracking whether gains follow the intended save-retrieve-update pathway.
- Key Findings: Across 7 base models and 4 agent frameworks, retained experience improves performance, but gains are highly capability-, model-, and framework-dependent. Agents with similar headline gains often rely on markedly different persistence pathways.
- Optimization: Introduces Hermes+, which applies 5 targeted interventions across the agent loop, effectively improving both the persistence gap and mechanism-evidence scores.
Related event: Princeton Introduces PAST-Bench for Evaluating AI Self-Improvement(2 posts)→
More from coding & agent
- New MCP Tool Enables Natural Language Domain Management via Namecheap API — modelcontextprotocol · 2026-08-05
- Twinmotion MCP Connector Automates Architectural Rendering and Video Export — modelcontextprotocol · 2026-08-05
- deep-research-mcp Adds Support for Gemini and Other DR Backends — PMinervini · 2026-08-05
- Dark Recesses of Ambient Memory: A Landfill of Logs Is Not a Brain — bfrench · 2026-08-05
- Agent Hallucination Disaster: Auto-Reply Bot Sends Wrong Price to Client — Trusttive11 · 2026-08-05
- Open Source Tool code-review-graph: Saving Tokens for AI Coding via Code Graphs — PrajwalTomar_ · 2026-08-05