DerivAudit: 17-21% of long-term agent memories unsupported by interaction history
newyorkuniversity · hf · 2026-10-01
An NYU team introduces DerivAudit, a framework auditing whether an LLM agent's persistent memories are actually supported by the interaction history available at write time. Compression can introduce relations or event statuses history never established.
Key findings:
- Broader pre-write history recovers support for nearly 60% of memories that look unsupported from citations alone;
- Yet 17-21% remain unsupported after expansion;
- Evidence expansion alone doesn't fix admission: unsupported memories are still frequently admitted across verification models, and expansion worsens admission on two backbones.
The audit separates three coupled requirements: evidence scope, compositional validity, and admission reliability.
More from coding & agent
- Edward Kmett ships Turbo Haskell: a GraalVM JIT for GHC Core that can compile GHC itself — rickasaurus · 2026-10-01
- Lance Martin to add data-first walkthrough to Claude eval skill after Hamel Husain feedback — RLanceMartin · 2026-10-01
- Personal Agent VM configs compared: Grokbot spends big on CPU, RAM and disk — op7418 · 2026-10-01
- Agent Lens launches as agent-native observability tool with cheap production judges and OTEL support — _ScottCondron · 2026-10-01
- Abacus AI Lets You Deploy Coding Agents and Review Fixes From WhatsApp or iMessage — bindureddy · 2026-10-01
- 6 tips to slash token costs in Claude Code workflows (ultracode) — majidmanzarpour · 2026-10-01