乐于分享
好东西不私藏

另一种AI风险:把所有AI都当成一回事来管控

另一种AI风险:把所有AI都当成一回事来管控
欢迎关注药事云端,了解国内外药事法规和技术趋势
译文by Thomas Tang

过去三个月,我一直在强调:在GMP环境中,AI的问题不在于“是否采用”,而在于“如何控制”。
三月,我提出了这一观点;四月,我在一封FDA警告信中展示了其实例;五月,我将其凝练为一句话:模型不会承担质量部门的责任——公司才会。
本月,我想将这一论点反转过来,因为第二种失败模式正悄然变得愈发普遍。
回顾过去30天的监管动态:USP发布了新标准和药品短缺报告;FDA发布了细胞与基因治疗指南,同时推进了新药批准、标签变更和警告信;英美药监局(MHRA–FDA)启动了联络项目;MHRA发布了关于医疗领域AI应用的证据基础;FDA还将首个计算机模拟工具纳入了ISTAND(新药创新科学与技术方法)试点项目。这样的一个月,并不罕见。
没有任何一个质量组织能对所有这些事项都施加最高级别的管控流程。然而,许多企业如今却对AI采取了一种条件反射式的做法——无论AI用于何种场景:生成GMP会议纪要、起草SOP初稿,还是输入批放行数据——都套用同样繁重的验证、同样详尽无遗的文档要求、同样冗长的审批链条。这种本能看似安全,实则并非治理,而是一种与风险脱节的摩擦,它正在训练员工将控制视为一场“表演”。
这里有个令人不适的事实:控制不足与控制过度,本质上是同一种病。二者都源于未能以“预期用途”和“风险”为锚点。四月Purolea公司的警告信就是风险控制不足的典型案例——AI生成的内容未经审核就直接进入记录系统。而当前的“过度控制反射”则是其反面:风险未加区分,将所有AI输出都当作批放行决策来对待。其实,ICH Q9早已提供了对症良方,它一直就在我们的工具箱里,只是被忽视了。
值得注意的是,监管机构在实施“相称性原则”方面,反而比许多制药企业做得更好。当FDA将那个由AI驱动的数字肝脏模型纳入ISTAND试点时,它并非笼统地认可“AI用于毒理学”,而是接受了一份意向书——这是三步资格认定中的第一步,且仅针对一个明确定义的使用场景。该工具必须围绕特定目的证明自身价值,而非凭借泛化的“能力”获得通行证。这正是将基于风险的相称性治理写入了监管路径本身。

三大启示

1. 按后果治理,而非按工具数量治理
一个用于起草会议纪要的AI,与一个支持批放行决策的AI,虽同属同一药品质量体系(PQS),但所需的控制深度应截然不同。若投入的努力不与潜在后果挂钩,便是双重浪费:既浪费成本,也损耗可信度。
2. 无差别的控制自有其代价。
无处不在的繁文缛节会挤占真正高GMP风险环节的关注资源,还会让审核人员养成“机械签字”的习惯。一种人人都执行却无人真正相信的控制,比没有控制更糟——因为它看起来合规,实则毫无防御力。
3. 相称性才是跨监管辖区的可扩展之道
USP、FDA、EMA、MHRA的监管产出正同步加速,且远未完全协调一致。你不可能为每个监管机构都维持一套独立的“最大强度”管控姿态。唯有在PQS内部建立基于风险、锚定预期用途的治理框架,才能在要求激增而共识滞后的情况下保持体系的一致性与韧性。

结语视角

过去四个月的核心教训,并非“加强控制”,而是“正确控制”。薄弱的治理会在规模化时暴露无遗,但若治理机制连“会议纪要”和“批放行决策”都无法区分,它终将因自身重量而崩塌,并拖垮团队对整个质量体系的信任。
因此,在问“该对AI施加多严的控制?”之前,请先问:这个AI在做什么样的决策?如果它出错,代价是什么? 答案将决定所需的严谨程度。除此之外,其余皆是表演,或是疏忽。唯有那条可辩护的中间道路,才值得我们立足。

原文:

The Other AI Risk: Controlling It Like It's All the Same

For three months, I have argued that AI in GMP fails on control, not adoption. In March, I made the case. In April, I showed it in an FDA warning letter. In May, I landed it on a single sentence: the model will not carry quality unit responsibility; the company will.

This month, I want to turn the argument around, because a second failure mode is quietly becoming more common.

Look at the last 30 days of regulatory output: USP standards and a shortages report; FDA guidance on cell and gene therapy alongside approvals, labelling changes, and warning letters; an MHRA–FDA liaison program; an MHRA evidence base on AI in healthcare; and FDA’s first in-silico tool into the ISTAND program. That is not an unusual month.

No quality organization can apply maximum ceremony to it all. Yet, that is exactly what many organizations now do reflexively with AI, treating every AI touchpoint—a GMP meeting summary, a first-draft SOP, a batch-release input—with the same heavy validation, the same exhaustive documentation, and the same sign-off chain. The instinct feels safe. It is not governance. It is friction that does not track risk, and it trains people to treat controls as theatre.

Here is the uncomfortable part: under-control and over-control are the same disease. Both come from failing to anchor on intended use and risk. The Purolea warning letter (April) was risk under-applied—AI output entering records with no review. The over-control reflex is the opposite: risk undifferentiated and every output treated as if it were a release decision. ICH Q9 is the cure for both, and it has been sitting in the toolbox the whole time.

Notably, regulators are modelling the proportionate approach better than many manufacturers. When FDA accepted the AI-driven digital liver model into the Innovative Science and Technology Approaches for New Drugs (ISTAND) pilot, it did not bless "AI for toxicology." It accepted a letter of intent, the first of a three-step qualification, for a narrowly defined context of use. The tool has to earn its place against a specific purpose, not a general capability. That is risk-proportionate governance, written into the pathway itself.

Three Implications

Govern by consequence, not by tool count. An AI that drafts a meeting note and an AI that supports batch release sit under the same pharmaceutical quality system (PQS) umbrella and demand radically different control depth. Effort should track consequence or it is wasted twice: once in cost, once in credibility.

Undifferentiated control has a body count of its own. Heavy ceremony applied everywhere crowds out attention from the touchpoints that actually carry GMP risk, and it conditions reviewers to sign reflexively. A control everyone performs but no one believes is worse than no control, because it looks defensible and isn’t.

Proportionality is what scales across jurisdictions. Regulatory output is accelerating across USP, FDA, EMA, and MHRA simultaneously, and it is not neatly harmonized. You cannot maintain a separate maximal posture for every agency. Risk-based, intended use-anchored governance inside the PQS is the only way to stay coherent when the requirements multiply faster than they converge.

Closing Perspective

The lesson across four months is not "control more." It is "control right." Weak governance gets exposed at scale, but governance that cannot distinguish a summary from a release decision collapses under its own weight and takes the team's trust in the quality system with it.

So, before asking how tightly to control an AI, ask what the AI decides and what it would cost if it were wrong? The answer sets the rigor. Everything else is theatre or neglect. The defensible middle is the only place worth standing. 


真的不打算点一下“看”?