MEMOIR benchmark: 117 synthetic oncology patients to test AI clinical memory
jefrankle · x · 2026-09-24
Olira released MEMOIR, an open-source benchmark for evaluating longitudinal clinical memory. It includes 117 synthetic oncology patients, 65,377 clinical events, and 3,617 questions spanning factual retrieval, rare-event detection, episode reconstruction, and longitudinal reasoning.
They compared general-purpose memory systems (mem0, Graphiti, Letta) against a purpose-built patient representation across thousands of clinical questions grounded in oncology records, noting that clinical memory has a different structure from conversational memory.
More from Research
- New DAYJOB benchmark: best agents complete only ~25% of real knowledge work — echen · 2026-09-24
- OpenAI Releases MentalHealthBench, an Open Benchmark Built with 80+ Clinicians — OpenAI · 2026-09-24
- Paperena Benchmarks AI Scientists Across Full Research Cycles: Writing, Reviewing, Revising — yeewhye · 2026-09-24
- 31 open reproductions of Jev's decision model benchmarked across 37 suites — tvnmsk · 2026-09-24
- ACL'23 outstanding paper: discriminative LMs may generalize better than autoregressive models — ysu_nlp · 2026-09-24
- 10 agent reruns reached the right neighborhood, none reproduced the key observation — rohanpaul_ai · 2026-09-24