Every phrase from the policy doc, still wrong: a minimal case that breaks answer-checking evals

YakDue6710 · reddit · 2026-10-07

An agent-eval practitioner shares a minimal adversarial case that breaks "is this text in the source" style checking:

Policy: unopened items returnable within 30 days for a full refund; opened items within 14 days for store credit only; refunds process in 5-7 business days.

Customer: "I opened the headphones I bought 10 days ago. Can I get my money back?" The bot answers yes — citing the 30-day window and full refund to the original payment method. Every number and phrase exists in the policy; nothing is invented. But the bot applied the unopened-items rule to an opened product, promising cash when only store credit applies. Any source-grounding check passes it, and the final sentence is even true on its own.

The author asks whether your eval/guardrail flags this, how you'd score the last sentence as supported or not, and offers more cases (near-miss numbers, dropped exceptions, conflicting documents).

Original post →

More from coding & agent

coding & agent channel →