PACT benchmark: one sentence of pressure raises AI rule violations 65%; no model clears unsupervised bar

baseten · x · 2026-08-28

Trace AI Labs released PACT (Pressure-Applied Compliance Testing), a benchmark testing whether enterprise AI assistants keep following workplace rules when breaking them is the convenient option. It covers 48 scenarios across 12 regulated domains (HIPAA, hiring law, GDPR), 9 kinds of everyday pressure, 3,364 items, and 23 models. Baseten sponsored the benchmark.

Key findings:

GLM-5.3-Flash is the best-performing GLM model, ranking #7 overall; the two highest-scoring models are both open-weight. Seven scenarios map to violations courts and regulators have already punished (Moffatt v. Air Canada, Mobley v. Workday, NYC MyCity).

Original post →

More from Models

Models channel →