Evaluating Semantic Caching: How to Verify LLM Rewrites Without Ground Truth?

Reasonable_Royal_621 · reddit · 2026-08-19

The author, working on CacheVerifier, discusses the challenge of evaluating a verification strategy inspired by TweakLLM, where a cheap LLM rewrites cached answers instead of rejecting them. Unlike binary rejection, rewriting generates new text with no ground truth labels. Proposed solutions have flaws: using an LLM judge adds cost/noise, similarity metrics contradict the project's thesis, and crude recall metrics yield weak results. The author seeks advice on evaluating free-form rewrites without definitive labels.

Original post →

More from Infra

Infra channel →