The CIOMS WG XIV report is built around the seven principles outlined in Table 1. They are introduced in the following sections, and selected pharmacovigilance use cases from the report are included to illustrate aspects relevant to each principle (see Table 2). These examples are illustrative rather than exhaustive, and their inclusion should not be interpreted as endorsement by CIOMS WG XIV of any specific method or approach. Please note that several of the use cases are proofs-of-concept and not deployed for real-world use, at the time of this publication. Further details and elaboration are provided in the CIOMS WG XIV report [1].
Table 1 Guiding principles for AI in pharmacovigilance from CIOMS WG XIV (here in abbreviated form)Table 2 Use cases from CIOMS WG XIV report2.1 Risk-Based ApproachIssues that could negatively affect a pharmacovigilance system’s behaviour and results, including errors and biases, should be identified, prioritised, and managed in a way that is proportionate to their risk. A risk assessment should be performed for each AI system aimed at supporting pharmacovigilance processes and should form the basis for a risk-proportionate approach applied throughout the AI system’s lifecycle from development to routine use.
Integrating AI systems in pharmacovigilance processes may affect patient safety and public health, user trust and engagement, efficiency, and other aspects such as data privacy or, where applicable, intellectual property. When determining the level of risk related to the implementation of AI within a pharmacovigilance system, key considerations include the AI technology itself, the context of use, the likelihood of risks materialising, their detectability, and their potential impact. Knowing when and how to mitigate risks requires being able to detect issues. This is usually based on a testing plan with key performance indicators laid out during the development of the AI system. Risk mitigation approaches can be proactive and/or reactive and depend on the AI system, the pharmacovigilance task, including the level of human oversight, and the type of risk. The risk strategy should be reviewed and adapted as needed at regular pre-planned intervals, or whenever the AI system shows performance issues.
Use case
Aspects relevant to risk-based approach
Case processing
Assessment of potential impact of inaccuracies on patient safety and pharmacovigilance outcomes. Compliance with regulatory and data privacy requirements considered, to mitigate risks related to these areas. Continuous monitoring of LLM performance and robust user training among risk mitigation strategies [3]
Deduplication
Assessment of potential impact of missed or false-positive duplicate reports. Risks in case series analyses mitigated by implementation of tool as part of decision-support system with humans in control. Risks in disproportionality analysis and signal detection yet to be evaluated [5]
Information synthesis
Data privacy and security concerns identified as important risks. Mitigation with e.g. use of closed environment for training and testing, and development of data privacy and security protocols. No patient safety risk currently anticipated in routine use [1]
2.2 Human OversightAlthough human oversight by itself does not guarantee perfect outcomes, it supports optimisation of performance of AI systems deployed in pharmacovigilance and increases trustworthiness and accountability. The complexity, sensitivity, and variable nature of pharmacovigilance tasks and data call for robust and risk-based human oversight mechanisms. Depending on the scope, extent, and intensity of human intervention, possible governance mechanisms include human-in-the-loop, human-on-the-loop, and human-in-command (Table 3) [12,13]. When using AI systems, pharmacovigilance professionals should be aware of pitfalls such as automation bias or confirmation bias [14].
Table 3 Modalities of human oversight; adopted from [12]The report also discusses the implications of the ever-increasing AI integration in pharmacovigilance processes on the workforce. The expected role of pharmacovigilance professionals in human oversight, as well as in the governance, design, evaluation, and implementation of AI systems, calls for sensible training, change management, and readiness strategies by pharmacovigilance organisations [15].
Use case
Aspects relevant to human oversight
Deduplication
Human experts actively involved during development and validation of deduplication pipeline. Ability to confirm or modify AI output (i.e. reference case selection) and audit performance within a decision-support tool during routine use [5]
Signal detection efficiencies
Human review essential due to complexity and variability of cases. Key role of domain experts in validation and interpretation of AI-generated outputs, to ensure alignment with clinical knowledge (e.g. identification of alternative causes for adverse events flagged by AI) [9]
Information synthesis
Manual verification process during development of tool, allowing users to assess relevance and accuracy of outputs. Level of human oversight required to mitigate possible automation bias and known risks of generative AI, e.g. hallucinations, to be further considered upon move to production [1]
AI artificial intelligence 2.3 Validity and RobustnessEnsuring the validity and robustness of AI solutions is central for patient safety, building trust and achieving the best possible value for end-users. The variable quality and consistency of adverse event reports and the complex nature of the studied drug–event relationships may require more extensive efforts than in other domains [16]. To invest resources optimally, pharmacovigilance professionals and decision makers must learn to critically appraise and evaluate proposed AI systems – whether they support their in-house development or procure them from other organisations [17].
The intended use and deployment domain for AI solutions in pharmacovigilance should be clearly defined, and performance evaluation targeted to these, as far as possible [18]. AI systems, like humans, will typically not achieve perfect performance on complex or ambiguous tasks. Therefore, performance is typically assessed statistically for a sample of cases referred to as the test set. Ideally, test sets must be large and diverse enough to demonstrate adequate and generalisable performance, detect biases, and identify circumstances where models may underperform [19]. When relevant benchmark methods and/or test sets exist, performance should be evaluated against these; however, for many pharmacovigilance tasks, robust benchmarks are still lacking.
The report further highlights and discusses special considerations related to recognition of low-prevalence events and patterns that are of special interest in pharmacovigilance [17]. Qualitative review of individual AI model outputs is promoted as a complement to statistical performance evaluation, and subgrouping and sensitivity analyses are recommended to assess generalisability and fairness and equity. The special challenges of assessing the validity and robustness of non-deterministic systems such as generative AI models with unbounded output and of evaluating the collaborative performance of human-AI teams are also considered.
Use case
Aspects relevant to validity and robustness
Deduplication
Multi-phase analysis to evaluate and validate deduplication pipeline prior to deployment and routine use, including qualitative error analysis that identified three main types of false positives [4]
Translation
Evaluation of generative output with continuous learning and performance monitoring by recalculating BLEU scores after each model update [6]
Diagnosis
Five independent international study sites, with substantial diversity of patients and phenotypes. Algorithm validated on three external data sets obtained at different clinical locations across two countries and two imaging devices with varying acquisition parameters. Case evaluations completed by a large group of retina specialists at multiple institutions [11]
BLEU Bilingual Evaluation Understudy 2.4 TransparencyTransparency helps build trust, enabling individuals and organisations to inspect and scrutinise the design and performance of AI systems, and ensuring regulatory compliance. AI use should be disclosed to affected stakeholders, which may include pharmacovigilance professionals, patients, health professionals, and regulatory authorities, depending on the application.
Key aspects of an AI model that should be shared include its intended use, the nature and extent of any human–computer interaction, model architecture (and possibly parameters), the data based on which the AI model was trained and validated, and any known limitations of the system. Explainability (or interpretability) reflects a specific form of transparency where the general principles and logic by which an AI model operates and has arrived at a specific output can be understood by humans. This can enable stakeholders to contextualise and influence an AI system’s output and may support model selection and troubleshooting during development. However, explainability is not an absolute requirement, nor does it mean that a system is fit for purpose or trustworthy. Whereas simpler models may be inherently explainable, a degree of post hoc explainability can sometimes be achieved through a layer of techniques applied on top of a more opaque AI model. However, these require their own validation [20]. Use of externally hosted AI models or vendor-provided services may limit the degree with which model transparency can be offered.
Transparency regarding an AI model’s assessed performance provides a bridge between theoretical capability and practical utility. Key aspects of an AI model’s performance that should be disclosed to stakeholders include the nature, size, and diversity of the test sets and their alignment with the intended deployment domain. Statistical performance evaluation should use relevant metrics and operating points for intended use and be complemented with qualitative review of representative AI model outputs.
Use case
Aspects relevant to transparency
Deduplication
Use of an inherently explainable AI model and pipeline where a probabilistic record linkage model allows inspection of elements that contributed to a given report pair being predicted as duplicates or non-duplicates. Extra efforts made to increase model, pipeline, and performance transparency leading to increased trust in AI model by human assessors [4]
Causality assessment
Use of inherently explainable AI model and pipeline. Training set, model features, and reference set described in scientific publication [8]
SQL generation
Full Python code specifying prompt, LLM model, and version included in scientific publication, together with detailed result files that link individual natural language inputs with SQL outputs, enabling external inspection and re-analysis of performance evaluation [7]
AI artificial intelligence, LLM large language model, SQL Structured Query Language 2.5 Data PrivacyEnsuring data privacy in AI solutions is central to building public trust and ensuring regulatory compliance consistent with ethical principles for the protection of human research laid out in the Belmont report [21].
The sensitive nature of pharmacovigilance data, which relate to health and may include patient information, requires more intensive safeguarding efforts than many other fields. Pharmacovigilance professionals and decision makers should recognise that existing procedures may need to be re-evaluated as AI increases the potential for patient re-identification, especially when linking databases.
Data processing must align with legislation such as the General Data Protection Regulation (GDPR) in the European Union (EU) or the Health Insurance Portability and Accountability Act (HIPAA) in the United States. In developing and deploying AI systems, pharmacovigilance data may be exposed to a broader range of stakeholders, and proactive measures like data minimisation, anonymisation, and encryption should be considered; such safeguards can be central to translating the broader principle of data privacy into operational practice. Ideally, organisations should conduct data protection impact assessments prior to deployment to identify and mitigate risks inherent in data processing.
The report further highlights and discusses special considerations for use of generative LLMs [22] related to data persistence, the secondary use of data for model training, and that the deployment environment is a key consideration for risk management; utilising private clouds within institutional firewalls provides essential control over data and prompts, whereas publicly hosted generative LLMs significantly increase the risk for data leaks. The report considers the implementation of advanced technical safeguards, such as federated learning, differential privacy, and private cloud deployment.
Use case
Aspects relevant to data privacy
Case processing
Research carried out in private cloud with access restricted to project members. Personally identifiable information redacted prior to data extraction step [3]
Translation
Service established on a private cloud with restricted access. Personally identifiable information redacted before translation, and original documents only accessible to local teams [6]
Diagnosis
Patient data deidentified for retrospective study in which Institutional Review Board (IRB) approval was sought to ensure ethical data handling
2.6 Fairness and EquityFairness and equity highlight the importance of identifying and managing potential unfair bias associated with the development or use of an AI solution that may result in a negative impact on groups [23]. The main risks are to perpetuate bias present in training data and inadequate performance for underserved populations. The benefits of AI in pharmacovigilance should be equitable across all groups.
Assessing AI solutions for common biases [24], adequate data representation, and explainability are key mitigation measures. Human biases and inadequate data can produce underperforming models and result in discriminatory harm. When reference data used to develop and test AI solutions is inadequate and under-represents subpopulations, it may threaten fairness and equity. Within the pharmacovigilance space, limitations of adverse event reporting systems with variable reporting rates across regions and lack of representative data in subpopulations (e.g. paediatric populations) are well known. The data limitations for subpopulations directly contribute to hampering the ability to analyse data, capture patterns, and make relevant inference in these subpopulations.
The report refers to the importance of evaluating and documenting data representation. AI solution explainability may highlight explicit bias, whereas model and performance transparency may help identify areas where bias may be introduced and provide user insight into expected performance and limitations. These insights may reduce the likelihood of misinterpretation and improper use and can inform the degree of human oversight required based on the data profile, risk assessment, and type of AI model. Identification of potential risk areas is challenging but key to preventing bias, discrimination, and suboptimal model performance.
Use case
Aspects relevant to fairness and equity
Information synthesis
Notes that inherent limitations and biases exist within safety data which may manifest within generative LLMs perpetuating bias. Relies on pharmacovigilance professional awareness of safety data limitations to limit impact of bias [1]
Diagnosis
Acknowledges that data under-represented specific groups (e.g. Asians), who may display different findings, and recommends further assessment with more diverse data. Suggests that AI solution may especially benefit currently under-served settings where retinal specialists may not be available, such as under-resourced or under-represented locales [10]
AI artificial intelligence, LLM large language model 2.7 Governance and AccountabilityRobust governance and clear accountability help ensure that AI systems are used responsibly and ethically throughout the AI lifecycle and are compliant with regulations, while fostering trust and transparency among stakeholders. It is advisable to nominate a governance body that oversees the lifecycle management of the AI system and determines accountable persons for the respective lifecycle phases. There should be measures in place to intervene or even disable an AI system if necessary. AI systems deployed in critical PV processes must be integrated into an organisation’s business continuity plan to ensure that safety monitoring and regulatory compliance are maintained if a system is unavailable or its performance degrades.
To facilitate the application of the guiding principles in real-world pharmacovigilance activities, the CIOMS WG XIV report presents a governance framework grid for AI solutions in pharmacovigilance (Fig. 1). The grid is composed of five lifecycle phases of an AI system: an initial requirement specification phase where business units typically provide input, followed by development, pre-deployment, post-deployment, and routine use. It is designed as a structured guide to identify key considerations to address each of the principles throughout these five lifecycle phases of the AI system. During the lifecycle, it may be necessary to go back to a previous phase to address certain needs and discoveries.
Fig. 1
Excerpt from governance framework grid. Complete grid can be found in CIOMS WG XIV report [1]. AI artificial intelligence, CIOMS WG XIV Council for International Organizations of Medical Sciences Working Group XIV
Traceability and version control are crucial aspects of managing AI systems, particularly in a regulated field like pharmacovigilance. Transparency between the development team and the pharmacovigilance organisation is crucial to ensure efficiency and that the system is fit for purpose. From a regulatory perspective, AI components used to support pharmacovigilance activities should be appropriately documented, for example, in the Pharmacovigilance System Master File (PSMF), in jurisdictions such as the EU/European Economic Area (EEA), United Kingdom (UK), and other regions that require or follow PSMF-based requirements [25].
Use case
Aspects relevant to governance and accountability
Deduplication
System administrators control and monitor deduplication pipeline and use of its output. Clearly defined roles specified in development, pre-deployment, and post-deployment stages [5]
Translation
Regular lifecycle governance executed by vendor and available on request to MAH. Notes that while responsibility for execution of translation lies with vendor, ultimate accountability remains with MAH [6]
Signal detection efficiencies
Notes that during proof-of-concept, the accountability of the methodology remains with the developer and that if the methodology would be integrated into a production setting, accountability would transition to the human subject matter experts [9]
MAH Marketing Authorisation Holder
Comments (0)