MemGym benchmarks long-horizon memory for LLM agents across coding and web tasks
dhruv2038 · x · 2026-07-26
MemGym proposes a long-horizon memory benchmark for LLM agents
The paper argues that most existing memory benchmarks only test whether a chatbot remembers user preferences, which says little about real agents operating over long tasks. To close that gap, the authors introduce MemGym, a benchmark that evaluates memory in realistic agent settings such as deep research, coding, and computer use.
Key points:
- MemGym unifies existing agent gyms and in-house memory-grounded pipelines behind one memory-reasoning interface.
- It spans five evaluation tracks grouped into four regimes: tool-use dialogue, multi-turn deep research, coding, and computer use.
- The benchmark reports memory-isolated scores, separating memory quality from reasoning, retrieval, and tool-use ability so strategies can be compared more cleanly.
- The authors also build synthetic pipelines for MemGym-CodeQA and MemGym-DR, with length control and ablation checks.
- For coding environments, they train MemRM, a lightweight reward model based on Qwen3-1.7B fine-tuned with QLoRA, to score compression quality faster than full Docker rollouts.
More from coding & agent
- Built with Codex, a parameterized cat generator is now open source — ring_hyacinth · 2026-07-26
- LLM Zoomcamp demo traces a Pydantic AI agent into DuckDB as 24 tables — Al_Grigor · 2026-07-26
- Combined MCP server adds Redshift querying and S3 Markdown semantic search — modelcontextprotocol · 2026-07-26
- OneQAZ launches an MCP connector for live crypto and stock market intelligence — modelcontextprotocol · 2026-07-26
- CREAO schedules a July 30 panel on getting AI from demo to production — thetripathi58 · 2026-07-26
- An AI agent clears Slay the Spire 2 on Ascension 8 and opens its harness and trajectories — bdsqlsz · 2026-07-26