Every phrase from the policy doc, still wrong: a minimal case that breaks answer-checking evals
YakDue6710 · reddit · 2026-10-07
An agent-eval practitioner shares a minimal adversarial case that breaks "is this text in the source" style checking:
Policy: unopened items returnable within 30 days for a full refund; opened items within 14 days for store credit only; refunds process in 5-7 business days.
Customer: "I opened the headphones I bought 10 days ago. Can I get my money back?" The bot answers yes — citing the 30-day window and full refund to the original payment method. Every number and phrase exists in the policy; nothing is invented. But the bot applied the unopened-items rule to an opened product, promising cash when only store credit applies. Any source-grounding check passes it, and the final sentence is even true on its own.
The author asks whether your eval/guardrail flags this, how you'd score the last sentence as supported or not, and offers more cases (near-miss numbers, dropped exceptions, conflicting documents).
More from coding & agent
- GPT-6-luna Unlocks More Reasoning Tokens via API: ~18k Tokens Scores ~80.5% on Terminal-Bench — LysandreJik · 2026-10-07
- Encrypted prompt injection: one Copilot model leaked secrets in half the tests — Haunting_Ganache_850 · 2026-10-07
- Enterprise AI rolls out backward: chatbots are the finish line, not the start — shashib · 2026-10-07
- Gradio's ML Intern can now build Gradio apps from a single prompt — Gradio · 2026-10-07
- Bug Hunt Benchmark retest: GPT-6.1 Sol recovers, Muse still cheapest strong agent — PawelHuryn · 2026-10-07
- One Claude skill plus Scenario MCP automates the entire video-editing workflow — smtabatabaie · 2026-10-07