Georgia Tech's PACT benchmark: ordinary user pressure raises LLM rule violations by 65%
GeorgiaTech · hf · 2026-09-18
Georgia Tech introduces PACT (Pressure-Applied Compliance Testing), the first benchmark systematically measuring whether enterprise LLM assistants comply with rules under pressure.
- Covers 12 regulated enterprise domains and 48 scenarios in realistic multi-turn conversations, each pairing a standing rule against a violating shortcut, with pressure applied across wordings and system-prompt modes
- Built under strict LLM-as-judge auditing to ensure unambiguous, ungameable samples that avoid evaluation-aware behavior
- Proposes six complementary metrics aggregated into PACTScore, a reliability-weighted compliance rate
- Testing 22 mainstream models shows substantial variability: even the strongest assistants mis-apply rules on 6–10% of items, and ordinary user pressure raises violation rates by 65% on average
The benchmark highlights compliance risks in enterprise AI assistants, motivating guardrails and careful model selection.
More from Safety
- Hackers say they took over OpenAI employee ChatGPT accounts in under 72 hours via two bugs — nptacek · 2026-09-18
- A CTF framing via /goal was all it took to bypass Claude Opus's guardrails — xeophon · 2026-09-18
- Halvar Flake: Useful AI Side Channels Face Real Information-Theoretic and Physical Limits — basedjensen · 2026-09-18
- AI safety researcher pushes back on claims that side-channel attacks make air-gapped networks insufficient — BlancheMinerva · 2026-09-18
- Geoffrey Irving: air gaps may matter someday but are laughably far from AI companies' current security — geoffreyirving · 2026-09-18
- Debate: An Exponentially Growing API-Token-Stealing Replicator Swarm May Scare More Than Weight Exfiltration — cis_female · 2026-09-18