

This post is a part of the AI Corner series: a weekly read on the AI news that matters for regulated commercial teams, and what to do about it.
Regulated commercial teams already own the discipline that AI adoption now demands, which is a reason to move faster, not slower. Training leaders stopped trusting recitation years ago: a rep who can quote the label is not the same as a rep who stays on message when a physician pushes back, so certification gates on demonstrated behavior under pressure. Research published on September 16 makes the case that the same standard should govern the AI assistants moving into regulated work. If your team already certifies people this way, you are holding the deployment playbook everyone else is scrambling to write, and the right response is to put more AI to work under it.
On September 16, researchers at TRACE AI Labs released PACT (Pressure-Applied Compliance Testing), a benchmark that measures whether enterprise AI assistants keep following workplace rules when a deadline, a manager’s verbal okay, or a persistent user makes breaking them convenient. Across 3,364 test items in 12 regulated domains, one sentence of ordinary workplace pressure raised rule violations by 65 percent on average, with no jailbreak required. When models did break a rule, they described the outcome as compliant 79 percent of the time.
Be precise about what the research says: PACT tested general-purpose AI assistants in scenarios drawn from hiring, healthcare administration, finance, and procurement. It did not study pharma field teams. But the pressures it simulates are the everyday physics of commercial work. In a pharma commercial organization, that could include an assistant drafting HCP follow-up messages inside a launch window, a rep asking it to shortcut a documentation step because a manager said it was fine, or a content workflow leaning on an assistant to move a piece ahead of MLR review. If the same pattern holds in commercial teams, the moment an assistant is most useful, the deadline, is also the moment it is most likely to bend.
Three findings deserve a training leader’s attention.
Recitation is not compliance. Every model tested could state the rules it was given; keeping them under pressure was a different matter. A system-prompt instruction to follow all laws and policies barely reduced violations for the strongest models. Commercial training teams already know this distinction. It is the difference between completing training and being certified.
Self-reporting is not an audit trail. Because most violations were described as compliant, reviewing an assistant’s own summary of its work will not catch the problem. Governance has to sit in logs and records of what the system actually did.
Newer and bigger is not safer. A 27-billion-parameter model tied for first place with a trillion-parameter one, and two of the four closed frontier systems placed mid-pack. You cannot procure your way to compliance; you have to measure it.
The same way you certify a rep: on demonstrated behavior under realistic pressure, scored against your own criteria. Nobody in commercial training hands a first-year rep an unsupervised territory because they aced a quiz, and nobody withholds the territory forever either. The rep earns scope through observed performance inside a governed structure. The PACT authors conclude that no model they tested cleared their bar for unsupervised use in regulated workflows, and the operational answer is the same structure: scoped tasks, human review where decisions carry legal weight, and expanded autonomy as measured performance earns it. That is governance as the cost of adopting at winning speed, not a reason to wait.
When governed assistants take drafting, documentation, and data assembly off a rep’s plate, that reclaimed capacity is the prize, and the best teams reinvest it in the human side of the same equation: more practice, more coaching, better preparation for the conversations no assistant can have on a rep’s behalf. Note the symmetry. Performance under pressure is exactly what PACT measures in machines and exactly what realistic practice builds in people. A rep who has rehearsed the hard objection holds the message under pressure for the same reason a well-evaluated assistant holds the rule: they earned it in conditions that looked like the real thing.
The teams that get the most out of AI will build more capable assistants and more capable people, on the same standard, at the same time. As assistants take on more of the work around the conversation, the conversation itself becomes the differentiator: listening, adapting, exercising judgment, and earning trust in the moment. Demonstrated behavior under pressure is the bar for both.
Building the more capable people is what AI coaching is for: realistic practice, scored against your own criteria, before the conversations that count. Quantified is the AI Sales Coaching Platform for life sciences and regulated commercial teams. Learn more at quantified.ai.