AI-driven pharmacovigilance has spent the past two years chasing model accuracy, on the assumption that better performance would translate directly into lighter workloads. Jason Bryant, general manager of AI platforms at ArisGlobal, and Samuel Wallis, head of case processing at Bristol Myers Squibb, explain why that assumption has broken down, and propose a trust coefficient as the metric that could fix it.
Bristol Myers Squibb (BMS) has been running AI-driven adverse event processing for six months now, long enough for a pattern to emerge that will be familiar anywhere validated AI meets human judgement. Although the accuracy scores have been tracking upwards, the busy scientists who should by now be relying on the system keep checking its work anyway.
This is odd, considering the experience BMS has built with its automated extraction tool. This pulls the structured data out of an incoming case report ahead of any human review. Roughly 400,000 individual case safety reports pass through the company each year, and since June 2025 around 60% of these have gone through automated extraction, a share BMS expects to keep growing as confidence builds. Whatever efficiency BMS can recover from that expansion depends less on the model, though, than on how far the organisation is able to step back from a quality-control regime originally built for manual review.
While some of the rechecking may be defensible caution, the more likely explanation is ingrained habit. BMS’s challenge is how to garner enough trust in the system’s overall track record that teams stop double-checking routine cases as a matter of course.
Reading trust from behaviour, not sentiment
The answer isn’t merely to ask reviewers how confident they feel about AI-enabled decisions; this is about determining to what extent they are actively changing the way they work.
A trust coefficient could answer that question. It would need to weigh several things together, rather than scoring the model on its own: how reviewers actually behave, how far along the governance and validation work has come, how much risk sits behind a given task, and how ready the organisation genuinely is to depend on machine output. The point isn’t to try to arrive at a single fixed formula, but rather to devise a structured way of deciding when it’s safe to lean on a validated system’s output within a defined context of use.
From checking every step to watching the ones that matter
Agentic AI adds urgency to this. Where earlier AI-powered tools were built around streamlining one task at a time, agentic systems can now carry a case through several stages before anyone human gets involved. (BMS has already put one AI model in place to review what a second one produces, well before a case reaches a reviewer at all.)
Human judgement isn’t pushed out; it moves further down the chain, away from compiling data and checking fields and toward interpreting evidence and sitting with the cases that stay genuinely uncertain. Keeping people inside every step of the process is a different discipline from keeping them positioned to step in — which only needs to happen when the risk or the clinical stakes actually demand it. Accountability does not move with the workload: responsibility for patient safety stays with qualified professionals regardless of how much groundwork a system absorbs.
BMS redesigned the process before it added the tooling
BMS’s wider lesson is that AI should be introduced to a workflow that has already been simplified, not layered onto one still carrying years of accumulated hand-offs. The company modernised its core safety platform and stripped out redundant steps before automating, and continues to do so through a rolling three-year programme.
Encouragingly, regulators’ aspirations are in step with such moves: the EMA and FDA published 10 joint principles for AI across the medicines lifecycle in January 2026, calling for validation proportionate to intended use and ongoing performance monitoring,[1] while CIOMS reached a similar view in its December 2025 Working Group XIV report, arguing that governance has to be designed in from the outset rather than retrofitted.[2]
Soaring case volumes add to the impetus for change. VigiBase, the WHO’s global safety-reporting database, held over 40 million case records by late 2024 — some 70% of them logged within the previous 10 years alone[3]. And that curve keeps steepening as medicines grow more complex and AE volumes climb every year. Whether AI can process that volume is no longer the open question.
What remains unresolved is at what point a validated system has earned the right to fewer routine checks. Right now, that judgement is being made separately by the agencies that regulate these systems, the vendors that build them and the pharma companies that deploy them. A trust coefficient would matter less as a formula than as providing useful common ground: a shared basis for that judgement, rather than three parallel definitions of what trust in a safety system actually means.
This article builds on themes discussed in a recent life sciences industry podcast, which can be accessed in full here.
About the authors
Jason Bryant is general manager of AI platforms at ArisGlobal, based in London. Trained originally as a data science actuary, he has built a career spanning fintech and health-tech, with a focus on human-centred AI product design, and previously ran a digital incubator at AstraZeneca before moving into agentic AI for drug safety.
Samuel Wallis is head of case processing at Bristol Myers Squibb, based in London, where he oversees pharmacovigilance operations handling one of the industry’s largest ICSR volumes. His transformation work spans global safety operations, including migrating BMS’s global safety database to the cloud and rolling out AI-enabled case processing.
References
[1]European Medicines Agency and U.S. Food and Drug Administration, ‘EMA and FDA set common principles for AI in medicine development’, 14 January 2026. Available at: https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0
[2]Council for International Organizations of Medical Sciences (CIOMS), ‘Artificial Intelligence in Pharmacovigilance’, CIOMS Working Group XIV report, Geneva, December 2025. Available at: https://cioms.ch/working_groups/working-group-xiv-artificial-intelligence-in-pharmacovigilance/
[3]Brand JS, Gauffin O, Sartori D, Fusaroli M, Sköld H, Bergvall T, Sandberg L, Wallberg M, Hjelmström P, Norén GN. ‘VigiBase: Resource Profile Update with a Summary of Global Patterns and Trends in Adverse Event Reports for Medicines and Vaccines’. Drug Safety. 2026;49(6):613–629. DOI: 10.1007/s40264-025-01642-6

