Audit: 23% of "wrong" cache hits in a semantic caching benchmark were identical prompts
Reasonable_Royal_621 · reddit · 2026-10-02
The CacheVerifier team tested whether a small verifier beats a plain similarity threshold for semantic caching, evaluated on the public SemCacheLMArena and SemCacheSearchQueries benchmarks — then audited the labels.
Findings:
- 727 of 3,117 "wrong" near-duplicate hits on LMArena (23.3%) are byte-identical after lowercasing and stripping punctuation — one differed only by a leading space
- Blind hand-labeling found 88% of those "wrong" hits were perfectly safe to reuse
- Relative comparisons at the same error rate barely moved, but absolute error rates were wildly inflated: a 0.97 threshold's 5.2% error rate corrects to 0.95%
- One of their own results flipped: raising the skip-verifier cutoff to 0.99 looked like +5.18 points, but is -1.24 after fixing labels — the verifier was mostly rejecting duplicates the benchmark mislabeled as errors
Their advice: run a dumb identical-text check on your eval labels before tuning thresholds. Caveats: one annotator, small samples. Repo and erratum are public.
More from coding & agent
- Dev builds SCP Foundation survival game in Godot with Claude Opus 5.5 Max — imjustnewatai · 2026-10-02
- MockAgent: real-time AJV validation to catch agent tool-call hallucinations and retry loops — Zealousideal-Room775 · 2026-10-02
- Deedy's AI video workflow part 2: Claude Code, OpenRouter and emotional TTS setup — FinanceYF5 · 2026-10-02
- Dev turns personal bookmark library into an agent-accessible context source for Claude Code — dheeraj_iosdev · 2026-10-02
- Six Coding Agents, One Repo: Isolated Runs All Broke, Chatting Agents All Passed — jokiruiz · 2026-10-02
- Dev's AI Agent Astra Generates Game Assets via Pure Procedural Code, Not Imagen — Dimillian · 2026-10-02