Oxford & Tsinghua Launch Agent Memory Challenge for Unified LLM Memory Eval
JundeMorsenWu · x · 2026-08-02
Co-organized by over 20 universities including Oxford and Tsinghua, the Agent Memory Challenge is now open for submissions, aiming to provide a fair, unified evaluation for long-term memory systems in AI agents.
Key Design:
- Unified Protocol: Participants only implement add and search endpoints. Answer generation, judging, and scoring are handled centrally to eliminate setup variances.
- Benchmark Scope: Features Text and Coding tracks. The Text track includes 1,559 conversations (150M characters) and 5K questions. The Coding track uses 344 real-world software engineering tasks across 12 GitHub repos.
- Competency Model: Evaluates explicit fact recall, multi-hop reasoning, memory governance (updates/conflict resolution), personalization, and privacy.
Submissions close on August 7, with separate leaderboards for academic methods and commercial products.
More from coding & agent
- Test: AI Agent Runs Facebook Ads Autonomously for $1,500/Month, Beats Humans — PrajwalTomar_ · 2026-08-02
- Agensis Launches Desktop App for Human-Agent Team Collaboration — jasonkneen · 2026-08-02
- Integrating AI Agents into CI/CD: Automating the Software Factory — TejasKumar_ · 2026-08-02
- New ChatGPT macOS Workflow: Send App Window Screenshots Directly to AI — jxnlco · 2026-08-02
- MCPRadar: Open-Source Security Scanner and Leaderboard for MCP Servers — tatar-sh · 2026-08-02
- Bought a $5k Mac Studio for local LLMs, ended up running hundreds of subagents — EverydayAI_ · 2026-08-02