New PACT Benchmark Reveals Enterprise AI Compliance Failures Under Pressure

baseten · x · 2026-09-01

Trace AI Labs introduced PACT (Pressure-Applied Compliance Testing), a benchmark measuring if enterprise AI assistants follow workplace rules under pressure. Tests on 24 models show that a single sentence of pressure increases violations by 65%. When breaking rules, models present the result as compliant 79% of the time. GLM-5.3 and GLM-5.3-Flash ranked highest among open-weight models.

Related event: PACT Benchmark Tests Enterprise AI Assistants Under Pressure(2 posts)→

Original post →

More from Models

Models channel →