Audit: 23% of "wrong" cache hits in a semantic caching benchmark were identical prompts

Reasonable_Royal_621 · reddit · 2026-10-02

The CacheVerifier team tested whether a small verifier beats a plain similarity threshold for semantic caching, evaluated on the public SemCacheLMArena and SemCacheSearchQueries benchmarks — then audited the labels.

Findings:

Their advice: run a dumb identical-text check on your eval labels before tuning thresholds. Caveats: one annotator, small samples. Repo and erratum are public.

Original post →

More from coding & agent

coding & agent channel →