Same system scores 51.4 vs 75.8 depending on judge — are memory benchmarks trustworthy?
Efficient_Joke3384 · reddit · 2026-10-05
A Reddit user highlights how unreliable AI memory benchmarks can be:
- Judge-dependent scores: The same set of answers scored 51.4 under token overlap F1 but 75.8 under an LLM judge — a gap of over 24 points.
- Opaque methodology: Which model reads the retrieved memories, which model judges the answers, and how many runs were performed all matter, yet are rarely reported alongside the numbers.
- Vendor-reported noise: Many published numbers come from the vendors themselves, with no clarity on run counts or variance.
The poster asks how the community decides whether a benchmark number means anything, and whether a published score has ever actually changed which memory tool people picked — kicking off a substantive debate on memory benchmark methodology.
More from Models
- Fine-tuned Llama 3.1 8B for medical decisions hits 84% per-field accuracy, only 30-34% perfect rows — Forsaken_Cut8542 · 2026-10-05
- Ornith 1.5 35B-A3B hits 180 tok/s on dual 5070 Tis, 3x faster than Qwen 27B at same agent scores — Excellent-Issue-5956 · 2026-10-05
- ChatGPT to merge Chat and Work, drop model picker—users eye cancellation — Diamond_Mine0 · 2026-10-05
- AI Daily: Gemini free tier cut to Flash-Lite, Aleph Alpha open-sources 78B MoE Kolibri-1 — testingcatalog · 2026-10-05
- Mobilerun scores perfect 116/116 on Google's AndroidWorld, beating Artemis — SimplyAnnisa · 2026-10-05
- Weekend iOS app + pSEO project burned through full Opus quota plus ~$450 across 7 threads — iannuttall · 2026-10-05