97.6% per-turn relevance hides 7.4% memory failures across 64,000 turns
Swimming_You_6475 · reddit · 2026-10-11
Drawing on 12,800 five-turn travel sessions (64,000 turns) across 4 product teams, the author shows a blind spot in agent evals: per-turn relevance hits 97.6%, yet 7.4% of sessions (947) re-offer an itinerary the user already rejected — compressed conversation state drops the rejection before turn 4, and per-turn scoring never sees it.
Key discussion points:
- Session-level scoring can verify a rejection survives the whole session; trajectory scoring can catch the re-offer as it happens
- The hard part: scoring legitimate reconsideration (real preference reversal) without conflating it with lost conversation state
- The author asks the community how they distinguish genuine preference changes from multi-turn state loss
More from coding & agent
- GhidraMCP updated for latest Ghidra with headless support: load plugin and start reversing — lauriewired · 2026-10-11
- Harvard Med workshop teaches building AI co-scientists with ToolUniverse's 2,700+ tools — marinkazitnik · 2026-10-11
- Delvetown agents keep producing impressive artifacts as AI town experiment evolves — lfschiavo · 2026-10-11
- Nous Research launches stealth coding and agentic reasoning model, free for limited time — Teknium · 2026-10-11
- Cloud agent users mock devs still lugging around laptops in viral meme — steipete · 2026-10-11
- MCP server plus a World of Warcraft environment: AI agents entering Azeroth? — djcows · 2026-10-11