Red-Teaming Semantic Cache Verifiers: 84% of Adversarial Pairs Slip Through
Reasonable_Royal_621 · reddit · 2026-09-04
After a semantic cache served a "pause subscription" answer to a cancel request at 0.87 similarity, the author built CacheVerifier and red-teamed it with DeepSeek-generated adversarial pairs on five failure axes (negation, action-word swap, direction reversal, entity swap, quantity swap), keeping only pairs clearing the 0.80 similarity gate (306/400).
Key findings:
- An off-the-shelf cross-encoder approved 84% of adversarial pairs (CI 80–88%); negation and quantity swaps were 96% blind spots. One confidently wrong approval scored 11.32 against a threshold of 0.
- Single-round testing was uselessly noisy (58%→85% swings); four rounds plus bootstrapping were needed.
- Domain fine-tuning on natural gray-zone data bought nothing: 88% on adversarial set, CIs overlapping with 84%.
- Mixing 3.8% adversarial data (446 rows from 223 held-out triples) into 11,271 natural rows dropped false approvals to 54% (CI 48–59%) while natural-data AUC stayed at 0.875.
More from Infra
- AMD's Threadripper Halo Station packs 96 cores and 576GB of HBM3E — ccerrato147 · 2026-09-04
- AeroJEPA fluid foundation model joins NVIDIA's PhysicsNeMo ecosystem — ricardovinuesa · 2026-09-04
- Building a €2-2.5k local AI rig for legal RAG and agentic coding: hardware picks debated — whatyathinkk · 2026-09-04
- Dual 3090 owners debate adding more cards: bigger local models vs parallel instances — Blues520 · 2026-09-04
- NousResearch brings one-click local model setup to Hermes Agent on NVIDIA systems — lifebypixels · 2026-09-04
- Regulated-industry dev seeks AI Gateway with Okta SSO and runtime policy enforcement — IrrepressibleInk · 2026-09-04