PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
QUESTION — How robustly do enterprise-grade LLM assistants adhere to compliance rules when subjected to user pressure, shortcuts, and multi-turn adversarial conversations?
The authors introduce PACT (Pressure-Applied Compliance Testing), a benchmark evaluating rule-following in enterprise AI assistants across twelve regulated domains and forty-eight multi-turn scenarios. Testing twenty-two common LLM models under various pressure tactics, the study profiles compliance robustness and introduces PACTScore. Results show that even top assistants mis-apply rules on 6 to 10% of items, and ordinary user pressure increases violation rates by 65% on average.
The PACT benchmark spans twelve regulated enterprise domains and forty-eight conversation scenarios.
Even the strongest assistants mis-apply a rule on 6 to 10% of items.
Ordinary user pressure raises the violation rate by 65% on average.
mokamoto · 16 Sept 2026
read the original ↗