Building a long-term memory benchmark for agents: what to add?
True_Mongoose_7073 · reddit · 2026-09-01
Developer is building a benchmark for agent long-term memory, currently featuring months of chat history, multilingual data, images, answer changes, and token usage tracking. Seeking feedback on missing features, especially edge cases encountered in real-world memory usage.
More from coding & agent
- Using Grok Bot to build college admissions dataset pipeline — lennysan · 2026-09-01
- Idea: 'Money Leak Hunter' Grok Bot for finance audit — lennysan · 2026-09-01
- From Discord Bots to a Multiplayer Agent Workspace — steipete · 2026-09-01
- OpenClaw puts the agent on your machine: 933 volunteers vs rented AI assistants — heyneighbor · 2026-09-01
- Dev shares a dirt-cheap approach to visual diffs — zeeg · 2026-09-01
- Grok Bot automates Shopify updates and supplier coordination — billyjhowell · 2026-09-01