Princeton's PAST-Bench Tests If Personal Agents Actually Improve From Accumulated Experience
princetonu · hf · 2026-08-05
A Princeton team introduced PAST-Bench, a benchmark designed to evaluate whether personal AI agents can translate retained experience (preferences, histories, skills) into better future performance.
- Design: Spans 26 scenarios and 204 episodes, toggling retained experience on and off under matched conditions.
- Findings: Across 7 base models and 4 frameworks, improvement is real but uneven. Agents with similar headline gains often differ in whether those gains follow the intended save-retrieve-update pathway.
- Solution: The team proposed Hermes+, applying 5 targeted interventions across the agent loop. It raises average gains and provides clearer pathway evidence, particularly on tasks requiring outdated state replacement.
More from coding & agent
- Genomic Intelligence models now callable from Claude, Cursor via MCP for gene prediction — julia_kiseleva · 2026-08-05
- Shanghai Disney MCP Server: Enabling LLMs to Query Real-Time Ticket Pricing — modelcontextprotocol · 2026-08-05
- ScanBIM MCP Released: Supports 50+ 3D Formats and Clash Detection — modelcontextprotocol · 2026-08-05
- Post-Compact Reminder: A Claude Code Hook to Prevent Rule Amnesia After Compaction — doodlestein · 2026-08-05
- OpenConfer: Open-Source Voice Infrastructure for Agents to Call Humans for Decisions — RichardsonDx · 2026-08-05
- Grok 4.5 + Blender MCP: Build 3D Scenes via Natural Language — elonmusk · 2026-08-05