Augmenting Medical Judgment in Signal Confirmation: Design Considerations for an AI-Enabled Causality Assessment Framework

Translating a causality framework into an AI-enabled system requires a deliberate design workflow. In this article, we propose such a system in the form of the Signal Confirmation through Omnichannel Pharmacovigilance Evidence (PV-SCOPE) approach. The approach is intended to reflect how causality assessment is conducted in practice while remaining computationally tractable and adaptable. Figure 2 illustrates the role of automation, AI, and ML across the signal evaluation continuum. It conceptually highlights how such techniques could facilitate access to the most informative data sources, enable holistic multimodal analysis, and support visualization for clinical insight and transparent decision-making.

Fig 2.Fig 2.

Conceptual role of automation, AI, and ML across the signal evaluation continuum. AI artificial intelligence, ML machine learning, PV pharmacovigilance, EHR electronic health record, RWD real-world data

Transparency would be increased because the system would be designed to expose the source of each evidence element, the direction of evidence, the treatment of missing or conflicting information, and the rationale presented to the human assessor. The model would not attempt to harmonize expert judgment, algorithms, and probabilistic reasoning by collapsing them into one score. Instead, it places these complementary modes of reasoning into a common evidence architecture so that their respective contributions and limitations can be reviewed explicitly.

3.1 Design Considerations

An AI-enabled causality assessment platform should be designed for routine signal confirmation while retaining sufficient flexibility to support rare, unforeseen, and high-impact safety issues, including potential black swan events [13]. Its design should align with CIOMS XIV as well as FDA and EMA principles for good AI practice, which emphasize efficiency, transparency, governance, human accountability, and appropriate validation [14,15,16]. The recent CIOMS XIV report is especially relevant because it explicitly identifies causality assessment of adverse drug reactions as an AI use case in pharmacovigilance and frames the role of AI as supporting expert medical review.

The PV-SCOPE paradigm would use interoperable ontological frameworks [17] and knowledge graphs [18] to support semantic alignment across heterogeneous data sources that are traditionally siloed and described using different terminologies, structures, and levels of abstraction. In this context, ontologies provide formal, explicit representations of pharmacovigilance-relevant concepts and their relationships, enabling consistent interpretation of safety data across evidence sources. For example, clinical trial data, observational real-world data, and nonclinical toxicology findings may describe related biological phenomena using different vocabularies or contextual assumptions [19, 20]. Absent semantic harmonization, such differences can obscure relationships and hinder coherent evidence integration during causality assessment. Shared ontologies therefore provide a common conceptual backbone by enabling consistent reference representations of key entities such as medicinal products, adverse events, biological pathways, and clinical phenotypes.

Knowledge graphs would build on these ontologies by explicitly encoding relationships between entities, together with contextual metadata, enabling evidence from diverse sources to be connected and interpreted within a unified semantic framework. Within PV-SCOPE, this semantic structure would support both the system and the human assessor in recognizing when disparate observations across evidence streams refer to related concepts, even when expressed differently at the source level. For example, a clinical trial reporting “elevated liver enzymes,” an observational study capturing “drug-induced liver injury,” and a nonclinical study describing “hepatocellular necrosis” may all reflect related hepatic injury processes. Semantic alignment would allow such observations to be considered jointly during causality assessment.

Performance gains should be established during controlled testing and continuously evaluated in production through predefined metrics that reflect end-to-end signal confirmation quality. Responsible deployment also requires predefined guardrails to prevent high-risk or unacceptable outputs, including never events [21]. The system should also fit within established pharmacovigilance governance, documentation, and accountability structures rather than assuming replacement of existing workflows.

3.2 Phases of Model Development

To reflect its intended function, PV-SCOPE should be conceptualized as a modular, phased system that mirrors the sequential, yet iterative, reasoning process used by experienced safety assessors. Figure 3 conceptually depicts the PV-SCOPE paradigm as an end-to-end signal confirmation framework. The figure outlines the proposed flow for building such a system, from data integration across multiple evidence sources through source-specific data preparation, model training, evaluation, evidence integration, and iterative refinement.

Fig. 3Fig. 3

Conceptual PV-SCOPE development flow. DB database, EHR electronic health record, FAERS FDA Adverse Event Reporting System, FDA Food and Drug Administration, ICSR individual case safety report, LLM large language model, PV-SCOPE Signal Confirmation through Omnichannel Pharmacovigilance Evidence, RWE real-world evidence, WHO World Health Organization

Operationally, the PV-SCOPE paradigm would serve to: (1) define the adverse event and case concept under evaluation; (2) identify and retrieve relevant evidence from source-specific modules; (3) extract and structure key evidentiary elements, including diagnostic certainty, temporality, confounding, biological plausibility, study findings, background rates, and strength of association; (4) display the direction, strength, uncertainty, and gaps in each evidence stream for human review; and (5) support a documented assessor conclusion, including rationale, residual uncertainty, and recommended next steps. While assessors require concise, signal-specific summaries and explicit uncertainty cues, governance stakeholders may require aggregate views to monitor performance trends, potential bias, and system behavior across products, patient populations, and time [22].

3.2.1 Data Integration Across Multiple Evidence Sources

The first phase should focus on systematic identification and characterization of relevant evidence sources. Each source contributes distinct evidentiary value [2, 23] and operates at distinct levels of inference (individual versus population). Given differences in structure, content, and interpretive requirements, a key consideration is to treat evidence sources as separate analytic domains rather than forcing early aggregation, thereby preserving source-specific context and uncertainty. Real-world data merit particular emphasis among these sources. Linked claims, electronic health records, and registry data can supply denominators, background incidence, longitudinal follow-up, and outcome ascertainment that spontaneous reports lack. Used appropriately, they enrich otherwise sparse evidence and strengthen population-level inference, provided their provenance and fitness for purpose are established.

This phase is intended to establish the evidentiary foundation upon which all subsequent analytic steps depend. There is evidence that generative AI approaches might facilitate faster access to some data sources [24], especially given the disparate data structures used in healthcare systems, particularly across different geographies [25, 26].

3.2.2 Source-Specific Data Preparation

The second phase should involve development of data preprocessing and feature engineering pipelines tailored to each evidence domain in preparation for data ingestion. Given the variability and incompleteness typical of pharmacovigilance data, preprocessing should address issues such as missingness, duplication, coding variability, and noise. In addition, downstream modeling considerations should account for class imbalance, which is common in safety datasets where true events of interest are relatively rare. Feature engineering can be applied to extract the most informative elements for causality assessment, such as time-to-onset patterns, dechallenge and rechallenge indicators, co-medications, diagnostic certainty markers, and contextual modifiers. A full list of pertinent elements can be found elsewhere [3].

For unstructured data, NLP-based approaches could be applied to convert free-text narratives into structured information. Importantly, preprocessing strategies should be adaptable to centralized, federated, or distributed data architectures, where only summary-level or transformed data may be available for analysis. In such settings, AI-enabled quality assurance mechanisms could facilitate automated detection of inconsistencies or implausible patterns and trigger targeted queries to data assessors.

3.2.3 Data Analysis and Interpretation

In this phase, source-specific analysis must be able to separate genuine evidentiary patterns from artifacts of data generation. An apparent time-to-onset distribution, for instance, may reflect a true pharmacological relationship, or it may simply mirror nonrandom reporting or collection behavior, such as stimulated reporting after publicity. Such patterns should be treated as hypotheses to be tested against their provenance rather than accepted at face value, and the system should flag when an observed pattern is plausibly an artifact of how the data arose. Analytic outputs should not be treated as final causality determinations, but as structured inputs into a subsequent higher-level integration process. In this context, AE severity and seriousness lie outside the weighting of causal evidence. They remain essential in the bigger picture, but they inform the benefit-risk and action thresholds applied after causality is characterized, rather than the causal assessment itself. The present design is scoped to causality characterization and treats prioritization and action as downstream functions. Importantly, system outputs should explicitly represent disagreement, ambiguity, and residual uncertainty as valid and informative states rather than as deficiencies, avoiding false impressions of certainty driven by incomplete or conflicting evidence. Such modularity enhances flexibility, enables targeted refinement of analytic components, and supports transparent inspection of how conclusions are derived from each evidence stream.

To facilitate efficient human-in-the-loop review, the platform should include a centralized dashboard summarizing preliminary assessments on the strength and direction of evidence from each stream. This dashboard could be designed to allow assessors to rapidly evaluate how compelling the evidence is and to identify gaps requiring focused review. By making ambiguity and disagreement visible, the system enables assessors to concentrate their effort where it is most needed. The assessor should be able to accept or reject system-suggested characterizations, document alternative interpretations, and record the final rationale independently of the AI-generated recommendation. Audit trails of data used in decisions and the decisions themselves will be particularly important given the opacity of some AI algorithms, as highlighted in early surveillance tool systems [27].

3.2.4 Model Training and Evaluation

Model training and evaluation should leverage curated historical pharmacovigilance data, including safety signals that have undergone prior organizational review, spanning both confirmed safety signals and unconfirmed or refuted signals. These signals are intended to serve as positive and negative controls, enabling supervised learning approaches to calibrate the system’s ability to distinguish causal from noncausal associations under real-world conditions. In traditional supervised learning settings, such data may be used directly for model training. In contrast, for systems based on LLMs and retrieval-augmented approaches, historical signals are more commonly used for evaluation, benchmarking, and iterative system refinement, rather than direct parameter training. In these contexts, the focus shifts to assessing the quality of evidence retrieval, reasoning, and synthesis. However, when historical signals are used for training or evaluation, model developers should guard against information leakage, particularly when prior assessments may have been included in public regulatory documents, labels, or literature.

Model evaluation should extend beyond conventional accuracy metrics because signal causality assessment is not a simple binary classification task. Performance should be assessed on the system’s ability to handle incomplete or conflicting evidence, preserve source-specific context, generate clinically plausible outputs, and align with expert judgment and regulatory outcomes. At minimum, AI-supported assessment should perform no worse than human assessment alone [12], and preferably should improve consistency, efficiency, or decision coherence. Because variability across assessors and data sources is a known limitation of current practice, validation should explicitly evaluate reproducibility, including repeated-run stability, inter-assessor concordance, within-assessor reproducibility, and comparison of human-alone versus AI-supported assessments.

Evaluation should also include stress-testing across different signal archetypes, such as rare serious events, class effects, signals dominated by spontaneous reports, or signals driven primarily by observational data [28]. In addition, validation should assess the full socio-technical system, including how assessors interpret AI outputs, whether uncertainty cues are understood correctly, and whether the interface influences confidence or decision behavior in unintended ways. Usability testing should therefore be incorporated before operational use.

3.2.5 Iterative Refinement

Finally, as part of a continual learning system, feedback loops and postdeployment performance monitoring become essential considerations. This would allow structured learning from real-world use, including assessor interactions, decision outcomes, and regulatory feedback. Monitoring should also include drift in data sources, changes in coding practices, shifts in reporting patterns, and assessor override patterns. This will enable recalibration of evidence integration and identification of systematic blind spots as part of an evolving pharmacovigilance ecosystem [13].

Comments (0)

No login
gif