Two AI guardrail vendors disagreed on 11 of 40 identical inputs — and no labeled dataset exists

WolfShoddy7443 · reddit · 2026-10-07

An independent security consultant ran 40 identical inputs through two guardrail systems on the same agent — one hosted, one local — and the two disagreed on 11 cases with neither obviously wrong. Both vendors offered confident but mutually contradictory explanations, and a third system used as tiebreaker disagreed with both on six of those cases.

The consultant's takeaway: there is no ground truth for guardrail decisions anywhere, and no labeled dataset exists to arbitrate them — a real gap in AI security practice.

Original post →

More from Safety

Safety channel →