Glasshouse v0.1 launches: an open memory benchmark with 2,847 questions over a 1.97M-token conversation
True_Mongoose_7073 · reddit · 2026-09-22
Frustrated that vendor memory benchmark numbers never reproduce and swapping the grading model shifts results more than the gap between systems, a developer released Glasshouse v0.1 as an open benchmark.
Key design points:
- Scale: 2,847 questions embedded in a conversation up to 1.97M tokens, in 10 languages with 50 photos; four conversation sizes from 1,882 to 103,572 turns, with the answer-bearing portion identical across all sizes, so scores reveal where systems break down as history grows.
- Per-axis reporting, no headline number: each dimension is scored separately, since a system can excel on one axis and fail another.
- Beyond plain recall: confidently returning an outdated fact scores worse than saying "I don't know"; flagging conflicting stored facts is the correct answer; the benchmark also checks whether systems admit something was never said.
The submission board is empty and the author has not submitted anything, so community testing via PRs/issues is welcome; company submission instructions are in the repo.
More from Research
- Google's ScientistTwo solves 80.4% of 107 top-venue ML problems autonomously — thisdudelikesAI · 2026-09-22
- Tencent Hunyuan's WebCraftBench tests web apps like software, matching human preference 85.3% — TencentHunyuan · 2026-09-22
- Sarah Hooker proposes $1 paper submission fee to stress-test auto research agents — MannyKayy · 2026-09-22
- OpenAI's new model reportedly solved 100+ open math problems in 24 days of training — BilelKort · 2026-09-22
- Boltzbit previews paper claiming BAST lets LLMs learn up to 1,000x faster than SOTA training — jmhernandez233 · 2026-09-22
- MiMo's GRS and GAR: scoring what makes an RL answer genuinely good, not just passing — tokenbender · 2026-09-22