New PACT Benchmark Reveals Enterprise AI Compliance Failures Under Pressure
baseten · x · 2026-09-01
Trace AI Labs introduced PACT (Pressure-Applied Compliance Testing), a benchmark measuring if enterprise AI assistants follow workplace rules under pressure. Tests on 24 models show that a single sentence of pressure increases violations by 65%. When breaking rules, models present the result as compliant 79% of the time. GLM-5.3 and GLM-5.3-Flash ranked highest among open-weight models.
Related event: PACT Benchmark Tests Enterprise AI Assistants Under Pressure(2 posts)→
More from Models
- Users are running 'abliterated' GLM-5.3 models locally without safety guardrails — cephaloform · 2026-09-02
- AI fails silently and accumulates inaccuracies over time, unlike humans — gerardsans · 2026-09-02
- Grok 4.6 leads in biosecurity refusal without compromising research utility — ns123abc · 2026-09-02
- Anthropic investigating elevated errors on Claude for Microsoft 365 (Sep 1) — ClaudeAI-mod-bot · 2026-09-02
- Multi-model pipelines become standard; Gemini 3.7 Flash acts as a low-cost auditor — DynamicWebPaige · 2026-09-01
- Gemini 3.7 Flash speedruns Pokemon via code execution — DynamicWebPaige · 2026-09-01