Prediction Models for Acute and Chronic Postoperative Pain in Patients with Cancer: A Systematic Review and Meta-Analysis

Introduction

Cancer continues to represent a significant public health challenge in the twenty-first century. Recent estimates from the International Agency for Research on Cancer indicate that approximately 20 million new cancer cases were diagnosed globally in 2022, accompanied by 9.7 million cancer-related deaths. Projections suggest that the annual incidence of cancer will surpass 35 million by 2050, reflecting a 77% increase primarily attributed to population aging and growth.1 Surgical resection remains the cornerstone of curative and palliative management for most solid tumors, and the majority of patients undergo at least one operation over the course of their disease.2 Despite advances in minimally invasive techniques and the worldwide adoption of Enhanced Recovery After Surgery (ERAS) protocols,3 postoperative pain remains a frequent, undertreated, and clinically distressing complication of oncologic surgery.4 In addition to its immediate burden, severe perioperative pain has been associated with stress-induced sympathetic activation and opioid-mediated immunomodulation. These mechanisms are known to reduce natural killer cell cytotoxicity and alter μ-opioid receptor signaling on tumor cells, potentially promoting the survival and dissemination of minimal residual disease during the perioperative period.5

Postoperative pain in cancer surgical populations comprises two clinically distinct yet interrelated entities. Acute postoperative pain (AOPP) develops within the first seven days after surgery, whereas chronic postsurgical pain (CPSP), as defined in the International Classification of Diseases, Eleventh Revision, refers to pain persisting beyond three months in the absence of recurrence, infection, or pre-existing pain.6 AOPP affects an estimated 55%–85% of patients undergoing oncologic surgery, with up to 40% reporting moderate-to-severe pain despite contemporary multimodal analgesia.4 The burden of CPSP varies markedly by procedure: a 2024 prospective cohort reported a three-month CPSP prevalence of 32.4% after thoracic surgery, with neuropathic features rising from 48.7% at three months to 71% at one year;7 the European multicenter NIT-1 survey documented six-month CPSP rates ranging from 6.9% after sternotomy to 16.2% after endometriosis surgery;8 and persistent pain after breast cancer treatment affects 25%–60% of survivors.9 Inadequately controlled pain is consistently associated with delayed functional recovery, prolonged hospital stay, increased pulmonary and cardiovascular complications, impaired quality of life, anxiety and depression, persistent opioid use, and elevated healthcare resource utilization.4,10

A characteristic aspect of postoperative pain following cancer surgery is the significant interindividual variability observed in pain trajectory, intensity, and chronicity among patients undergoing similar procedures. This heterogeneity reflects a multifactorial pathogenesis encompassing demographic, surgical, psychological, genetic, and neurobiological determinants.11 Concurrent advances in pharmacogenomics and integrative biomarker research, exemplified by the National Institutes of Health Acute to Chronic Pain Signatures program, have catalyzed a shift in perioperative pain management from reactive, protocol-based analgesia toward proactive, individualized risk stratification within the framework of precision perioperative medicine.12 Within this framework, the early identification of patients at heightened risk for severe acute postoperative pain (AOPP) or incident chronic postsurgical pain (CPSP) is crucial for implementing effective interventions. This proactive approach facilitates the deployment of preventive strategies, such as tailored multimodal analgesia, regional anesthesia, and structured psychological prehabilitation, prior to the establishment and persistence of pain.13

Reflecting these priorities, the past decade has seen a substantial proliferation of clinical prediction models for postoperative pain in patients with cancer, spanning methodologies from traditional logistic regression and nomogram-based visualization to advanced machine-learning approaches such as random forest, gradient boosting, and deep learning.14,15 Despite this expansion, the evidence base remains fragmented. A 2024 systematic review of seventeen AOPP prediction models found that all were rated at high risk of bias under the Prediction Model Risk of Bias Assessment Tool (PROBAST), and only three had undergone external validation.16 A parallel review of nineteen CPSP prediction models reported AUROC values from 0.658 to 0.816, again with all models at high risk of bias and a near-universal absence of external validation.17

To date, no review has comprehensively synthesized prediction models for postoperative pain exclusively in the adult oncologic surgical population. Three gaps motivate the present study. First, prior reviews have been confined to specific cancer subtypes or to a single pain phenotype, leaving the integrated landscape of AOPP and CPSP uncharacterized. Second, methodological appraisal in earlier reviews has been inconsistent, with few systematically applying PROBAST—the instrument designed specifically for prediction-model research.18 Third, no previous review has conducted a quantitative meta-analytic synthesis of predictive performance by integrating AUC, sensitivity, and specificity across training and validation cohorts using SROC curves within this population. To address these gaps, we conducted a systematic review and meta-analysis to identify, characterize, and quantitatively synthesize the clinical predictors and discriminative performance of prediction models for both AOPP and CPSP in adult patients undergoing cancer surgery, providing an evidence-informed foundation for the rational selection, refinement, and clinical deployment of perioperative pain risk-prediction tools in oncology.

Materials and Methods Search Strategy and Selection Criteria

The protocol was registered in PROSPERO on 29 April 2026 (registration No. CRD420261377343). Because registration occurred after the search end date (February 2026), the registration is not described as prospective. The review was conducted in accordance with the PRISMA 2020 statement,19 the PRISMA extension for Diagnostic Test Accuracy (PRISMA-DTA), PRISMA-S,20,21 and published guidance for preparing and conducting systematic reviews in regional anesthesia and pain medicine22,23. Protocol deviations: During revision, additional post hoc analyses were introduced to strengthen the methodological robustness and clinical interpretability of the review. These included separate analyses of AOPP and CPSP, multilevel random-effects meta-analysis accounting for clustering of multiple model estimates within studies, CR2/Satterthwaite cluster-robust inference, 95% prediction intervals, and cancer-specific analyses within each pain phenotype. These additional analyses were based on the existing extracted data and did not alter the eligibility criteria or study selection. The manuscript title was also refined to more accurately reflect the focus on prediction models for AOPP and CPSP; this editorial change did not alter the review question or overall scope.

A systematic and comprehensive search was performed across four core biomedical electronic databases—PubMed, Embase, Web of Science, and the Cochrane Library—from database inception to February 2026. The search strategy was developed within the PICOS framework (Participants, Intervention, Comparison, Outcomes, Study design; Table S1, Multimedia Appendix 1) and combined Medical Subject Headings (MeSH) and free-text terms (Table S2, Multimedia Appendix 1). To minimize the risk of missing eligible records, we additionally screened the reference lists of eligible studies, relevant guidelines, and prior reviews, and conducted forward citation tracking in Web of Science. No trial registries were searched.

Prediction research encompasses both diagnostic models, which estimate the probability of an existing condition, and prognostic models, which estimate the risk of future clinical outcomes.18,24 This review included all primary studies reporting the development and/or validation of prediction models, tools, or risk scores for postoperative pain in patients with cancer.

Studies were eligible if they (1) enrolled consecutive adult patients with pathologically confirmed cancer undergoing oncologic surgery, with postoperative pain (acute or chronic) as the predicted outcome, across all cancer types (including but not limited to breast, hepatic, and colorectal cancers); (2) developed and/or validated a prediction model, tool, or risk score for postoperative pain in this population; and (3) employed a retrospective, cross-sectional, or prospective design. No language restrictions were applied at screening; non-English records were translated by professional native-speaking translators and independently reviewed in duplicate by two team members to assess eligibility and data extractability. Studies were excluded if they were (1) duplicate publications; (2) reviews, case reports, editorials, or conference abstracts; (3) studies whose full text could not be retrieved despite author contact; or (4) studies that did not report valid or extractable data relevant to the prespecified review objectives.

Data Extraction

Duplicate records were identified and removed using EndNote 21 (Clarivate). Titles and abstracts were screened, followed by full-text assessment against the eligibility criteria. Data were extracted using a form based on CHARMS,25 including study identifiers, patient characteristics, candidate and final predictors, model-development methods, discrimination, calibration, and 2×2 performance data. When a study reported multiple models, all models with a usable AUC were retained, but their shared study population was accounted for statistically. For the available-validation analysis, internal validation was used when reported; otherwise external validation was used, so each model contributed only once. The exact study/model membership of every pooled estimate is provided in Table S3. Corresponding authors were contacted for unclear or incomplete data. Two reviewers (X.X. and H.N.) independently extracted and cross-checked the data; disagreements were resolved by discussion, with Y.L. adjudicating unresolved cases.

Quality Assessment

Two reviewers (X.X. and H.N.) independently applied the Prediction Model Risk of Bias Assessment Tool (PROBAST) to evaluate the risk of bias and concerns regarding applicability of each included study; a third senior reviewer (Y.L.) adjudicated unresolved disagreements. PROBAST assesses risk of bias across four domains—participants, predictors, outcome, and analysis—each rated as low, high, or unclear. Applicability was assessed across three corresponding domains—participants, predictors, and outcome—using the same three-level scheme.

Statistical Analysis

All analyses were performed in R 4.4.2. AUCs were analyzed with metafor and clubSandwich, diagnostic 2×2 data with mada, and figures with ggplot2. AUCs and their standard errors were transformed to the logit scale.26 Because the 53 models arose from only 29 studies, the primary AUC syntheses used multilevel random-effects meta-analysis with model estimates nested within studies. Heterogeneity variances were estimated by restricted maximum likelihood (REML), and confidence intervals used CR2 cluster-robust standard errors with Satterthwaite small-sample degrees of freedom at the study level. Results are reported with both 95% confidence intervals and 95% prediction intervals. For sensitivity and specificity, bivariate random-effects models were fitted to unique-study 2×2 data and SROC curves were constructed.

Training and available-validation performance were synthesized separately. AOPP and CPSP were treated as distinct primary clinical phenotypes. Cancer-specific analyses were then conducted within each pain phenotype; strata with fewer than three independent studies were shown descriptively and not pooled. The mixed AOPP/CPSP estimates were retained only as secondary summaries.

Heterogeneity was assessed with the Cochran Q test and summarized by the estimated between-study and within-study variance components and an approximate total I2.27 REML was preferred to the DerSimonian–Laird estimator because substantial heterogeneity was anticipated and DerSimonian–Laird can underestimate heterogeneity and produce overly narrow intervals. CR2/Satterthwaite inference was used for the dependent-model primary analyses. In the one-model-per-study sensitivity analysis, REML with the Hartung–Knapp adjustment was used because it incorporates uncertainty in the heterogeneity estimate and generally provides better small-sample interval coverage than normal-based DerSimonian–Laird inference.28

The primary subgroup variables were pain phenotype (AOPP or CPSP) and cancer type (breast, gastrointestinal, lung, or other), because these strata represent different clinical populations, prediction windows, and outcomes. Calibration reporting was synthesized in a structured study-level table (Table S4). Calibration was not quantitatively pooled because no study reported a calibration slope or calibration-in-the-large estimate with sufficient uncertainty information for meta-analysis. Additional exploratory subgroup analyses examined predictor category, study design, data source, and internal validation method within each pain phenotype using the same clustered model; strata with fewer than three independent studies were presented descriptively and not pooled.

Two sensitivity analyses were conducted. First, one representative model per study—matching the models already selected in the authors’ original single-model Excel datasets—was synthesized by REML with the Hartung–Knapp adjustment. Second, all models belonging to one study were removed together in leave-one-study-out analyses. Exploratory small-study effects were assessed only in the independent one-model-per-study datasets using Begg and Egger tests and funnel plots.29 A two-sided P value < 0.05 was considered statistically significant.

Results Study Selection

A total of 6,104 records were identified through database searching, and 2 additional records through manual reference screening and citation tracking. After removal of 1,059 duplicates in EndNote, 5,047 unique records were screened by title and abstract, yielding 89 records for full-text review; 4 could not be retrieved despite contacting the corresponding authors. The remaining 85 full-text articles were assessed for eligibility, and 56 were excluded (39 with outcomes that did not align with the prespecified objectives, 14 with ineligible populations, 2 non-original publications, and 1 with an inappropriate study design). Ultimately, 29 studies reporting 53 prediction models were included.14,30–57 The selection process is summarized in Figure 1 (PRISMA 2020 flow diagram).

A flowchart of study identification and screening process for a review.

Figure 1 PRISMA 2020 flow diagram of study identification, screening, and inclusion.

Characteristics of Included Studies

The detailed characteristics are summarized in Table 1 and Table 2. Collectively, the 29 studies reported 53 prediction models for cancer-related postoperative pain, published between 2015 and 2026, with sample sizes ranging from 44 to 203,942 participants and a reported pain prevalence of 1.4% to 72.5%.

Table 1 Baseline and Methodological Characteristics of the 29 Included Studies

Table 2 Methodological and Clinical Characteristics of the 29 Included Studies

Twelve studies were prospective and 17 were retrospective. Most (25/29) were single-center, whereas 4 were multicenter. By target population, breast cancer and gastrointestinal cancer were each addressed in 10 studies, 5 focused on lung cancer, and the remaining 4 examined other cancer types. The numerical rating scale (NRS) was the most frequently used pain assessment instrument (16 studies), followed by the visual analogue scale (VAS; 6 studies) and the Douleur Neuropathique 4 (DN4) questionnaire (2 studies); the remaining studies adopted heterogeneous outcome definitions, including the VAS combined with the WHO analgesic ladder, diagnostic codes with long-term opioid use, the modified Brief Pain Inventory, the IASP pain classification, and a clinical criterion for phantom limb pain.

For missing-data handling, 1 study did not report its approach, 4 applied imputation, and 24 performed complete-case analysis. Only 2 studies retained continuous variables in their original form, whereas the other 27 categorized them before modeling. Nomograms were the most common presentation format (9 studies), followed by machine-learning performance tables (6 studies) and full regression equations only (5 studies); two studies presented conventional logistic regression tables, and the remainder used mixed formats.

Characteristics of Included Prediction Models

Across the 53 prediction models, logistic regression (LR) was the most frequently used algorithm (20 models), while the remaining 33 employed machine-learning techniques: random forest (n = 6), gradient boosting machines (n = 6), neural networks (n = 4), decision trees (n = 3), support vector machines (n = 2), naïve Bayes (n = 2), k-nearest neighbors (n = 1), and other less common approaches (n = 9). The AUC was reported across the models and ranged widely from 0.407 to 0.960 (Table S5, Multimedia Appendix 1). Sensitivity and specificity were additionally reported in 17 studies (37 models), with sensitivity ranging from 0.008 to 1.000 and specificity from 0.514 to 1.000.

Calibration reporting was heterogeneous and generally incomplete (Table S4). Twelve studies reported no calibration information. Of the 17 that did, 10 presented calibration plots without a directly poolable parameter, 4 combined a goodness-of-fit test with a plot, and 3 reported numerical calibration metrics without a plot. Selected numerical results included mean absolute calibration error, Hosmer–Lemeshow P values, an integrated calibration index, or a Brier score; however, no study reported a calibration slope or intercept with a standard error, so quantitative calibration pooling was not possible.

All 29 studies undertook some form of model validation: 22 performed internal validation only, whereas 7 (24.1%) also conducted external validation. Internal validation used cross-validation in 18 studies, random splitting in 6, and bootstrap resampling in 5.

Predictors

The candidate predictors spanned six conceptual categories: demographics and baseline characteristics; disease- and tumor-related characteristics; treatment-related factors; pain- and symptom-specific factors; psychosocial and behavioral factors; and examinations, investigations, and biomarkers. In total, 159 distinct predictors were extracted, with 3 to 11 variables incorporated per model. The five most frequently selected predictors were age, preoperative pain, radiotherapy, tumor size, and anxiety.

Risk of Bias and Applicability (PROBAST)

The risk of bias and applicability of the 29 included studies (53 models) were assessed using PROBAST (Figure 2 and Table S6, Multimedia Appendix 1). Overall, 8 studies (11 models) were judged at low risk of bias, 10 studies (19 models) at unclear risk, and 11 studies (23 models) at high risk. For applicability, 16 studies (21 models) raised low concern, 3 studies (12 models) unclear concern, and 10 studies (20 models) high concern.

A stacked bar graph showing study ratings for risk of bias and applicability domains and overall.

Figure 2 Risk of bias and applicability of the 29 included studies assessed with the Prediction Model Risk of Bias Assessment Tool (PROBAST). Bars show the proportion of studies rated as low, unclear, or high risk across the four risk-of-bias domains (participants, predictors, outcome, and analysis), the three applicability domains (participants, predictors, and outcome), and the overall judgements.

Abbreviation: ROB, risk of bias.

Domain-specific weaknesses were evident across the four risk-of-bias domains. In the Participants domain, 13 studies (44.8%) were at high risk, primarily owing to single-center designs and limited sample representativeness. In the Predictors domain, 9 studies (31.0%) were at high risk, mainly because predictor sets were not pre-specified, predisposing to overfitting. In the Outcome domain, 1 study (3.4%) was at high risk because outcome ascertainment relied on patient self-report and telephone follow-up, raising the potential for recall bias. In the Analysis domain, 7 studies (24.1%) were at high risk, predominantly owing to insufficient sample size relative to the number of candidate predictors. Applicability concerns were identified in 8 studies (27.6%; Participants), 5 studies (17.2%; Predictors), and 2 studies (6.9%; Outcome), reflecting overly restrictive eligibility criteria and reliance on predictors that are difficult to ascertain in routine practice.

Meta-Analysis of Discriminative Performance

Pain-phenotype-specific multilevel random-effects meta-analyses were the primary analyses. All models with usable variance were included, with multiple models clustered within their source study. Mixed-pain overall estimates were secondary.

Training phase. Seventeen AOPP models from 13 studies yielded a pooled AUC of 0.83 (95% CI 0.76–0.88; 95% PI 0.54–0.95; total I2 95.5%), whereas 14 CPSP models from 11 studies yielded 0.79 (0.75–0.83; PI 0.62–0.90; I2 94.6%) (Figure 3). In the bivariate diagnostic synthesis, AOPP sensitivity was 0.82 (0.74–0.88), specificity 0.83 (0.72–0.90), and SROC area 0.89 (6 studies); CPSP sensitivity was 0.71 (0.47–0.88), specificity 0.75 (0.65–0.82), and SROC area 0.79 (4 studies) (Figure 4).

A forest plot of area under the receiver operating characteristic curve for AOPP and CPSP models.

Figure 3 Model-level AUC estimates in the training phase, presented separately for AOPP and CPSP. Diamonds show multilevel REML pooled estimates with study-clustered CR2/Satterthwaite 95% confidence intervals; lighter lines on pooled rows show 95% prediction intervals.

Abbreviations: AUC, area under the receiver operating characteristic curve; CI, confidence interval; PI, prediction interval.

Notes: Right-side columns report AUC (95% CI) for each model and 95% PI for pooled rows.

Two plots showing SROC curves and pooled sensitivity and specificity for AOPP and CPSP.

Figure 4 Pain-phenotype-specific diagnostic performance in the training phase. (A) SROC curves and study operating points for AOPP and CPSP; (B) bivariate random-effects pooled sensitivity and specificity with 95% confidence intervals.

Abbreviations: SROC, summary receiver operating characteristic; AOPP, acute postoperative pain; CPSP, chronic postsurgical pain.

Notes: The right-side column reports each bivariate pooled estimate with its 95% CI.

Available-validation phase. Seventeen AOPP models from 13 studies yielded a pooled AUC of 0.80 (95% CI 0.76–0.83; 95% PI 0.66–0.89; total I2 82.3%), whereas 29 CPSP models from 16 studies yielded 0.75 (0.70–0.80; PI 0.50–0.91; I2 97.9%) (Figure S1, Multimedia Appendix 2). AOPP sensitivity was 0.75 (0.70–0.79), specificity 0.80 (0.66–0.89), and SROC area 0.79 (6 studies); CPSP sensitivity was 0.72 (0.59–0.82), specificity 0.72 (0.61–0.81), and SROC area 0.78 (9 studies) (Figure S2). The secondary mixed-pain pooled AUCs were 0.81 (0.78–0.84; PI 0.62–0.92) in training and 0.78 (0.74–0.81; PI 0.57–0.90) in validation.

Sensitivity Analysis

The one-model-per-study sensitivity estimates were close to the clustered primary results: training AOPP 0.83, training CPSP 0.79, validation AOPP 0.80, and validation CPSP 0.76. In leave-one-study-out analyses, pooled AUCs ranged from 0.81 to 0.84 for training AOPP, 0.78 to 0.80 for training CPSP, 0.79 to 0.81 for validation AOPP, and 0.74 to 0.77 for validation CPSP (Figure 5; Figure S3). Thus, no single study materially changed the phenotype-specific conclusions.

A mixed figure showing two leave one study out forest plots and two exploratory funnel plots for AUC.

Figure 5 Training-phase robustness analyses by pain phenotype. (A and B) Leave-one-study-out pooled AUC after all models from the named study were omitted; (C and D) exploratory funnel plots using one representative model per study.

Abbreviation: AUC, area under the receiver operating characteristic curve.

Notes: In panels (A and B), the right-side column reports the pooled AUC (95% CI) after each study was omitted.

Subgroup Analyses

Pain phenotype and cancer type were treated as the principal clinical strata (Table 3; Figure S4 and Figure S5). For AOPP, gastrointestinal-cancer models were the only cancer stratum with enough independent studies for pooling (training AUC 0.83, 95% CI 0.74–0.89; validation 0.80, 0.74–0.85). AOPP lung-cancer strata contained two studies and other cancers one study, so these were not pooled. For CPSP, breast-cancer models yielded AUCs of 0.76 (0.70–0.81) in training and 0.72 (0.63–0.79) in validation. Lung-cancer CPSP models yielded 0.85 (0.63–0.95) in training and 0.82 (0.63–0.93) in validation, but each estimate was based on only three studies and had a wide prediction interval. Other-cancer CPSP models were not pooled in training and yielded 0.80 (0.74–0.84) in validation (3 studies).

Table 3 Primary Pain-Phenotype and Cancer-Specific Meta-Analyses of Area Under the Receiver Operating Characteristic Curve

Additional phenotype-stratified exploratory analyses by predictor category, study design, data source, and internal validation method are shown in Figures S6–S12. Several strata contained fewer than three independent studies and were not pooled; pooled strata generally had wide confidence and prediction intervals. These results were treated as hypothesis-generating and did not alter the primary phenotype- and cancer-specific conclusions.

Small-Study Effects

Exploratory small-study-effect tests were conducted only in the one-model-per-study datasets. Egger’s test suggested asymmetry for training AOPP (P = 0.004), whereas the corresponding Begg test was not significant (P = 0.076). Neither test was significant for training CPSP, validation AOPP, or validation CPSP (all P ≥ 0.083). Given the small numbers of studies and substantial heterogeneity, these tests were interpreted cautiously and not as definitive evidence for or against publication bias (Figure 5 and Figure S3).

Discussion Principal Findings

This review identified 29 studies reporting 53 prediction models, but the clinically relevant results differed by pain phenotype. AOPP models showed pooled AUCs of 0.83 in training and 0.80 in validation; CPSP models showed 0.79 and 0.75, respectively. The wide prediction intervals—especially for CPSP validation—show that performance in a new setting may be substantially lower than the average. The mixed AOPP/CPSP pooled values were therefore retained only as secondary summaries. Among cancer strata, gastrointestinal-cancer AOPP models were relatively consistent, while lung-cancer CPSP models had higher average AUCs but were supported by only three studies and very wide intervals. These findings do not establish that any current model is ready for routine clinical use.

The included models have potential value for preoperative risk stratification, but the present evidence does not justify deployment. Age, preoperative pain, radiotherapy history, tumor size, and anxiety were frequently selected and are readily available in routine records; however, repeated selection does not demonstrate transportable predictive performance or clinical net benefit. These variables should be regarded as candidates for rigorous external validation rather than as a ready-to-use bedside score.

Among these five predictors, preoperative pain is the most consistent and biologically plausible determinant of postsurgical pain. A 2016 systematic review and meta-analysis confirmed preoperative pain as a significant risk factor for persistent pain after breast cancer surgery,58 an association underpinned by a clear neurobiological mechanism: heightened preoperative pain sensitivity—reflected, for example, in a lower pressure pain threshold—correlates with the intensity of acute postoperative pain.59 Among psychological predictors, anxiety stands out. Preoperative psychological symptoms, including anxiety, depression, and sleep disturbance, substantially increase the risk of CPSP and of unfavorable pain trajectories;60 a recent machine-learning study further suggested that specific anxiety-related symptoms may carry greater predictive power than generalized anxiety scores.61 Younger age has consistently been identified as a significant stratification variable, as evidenced by numerous studies. Notably, a 2016 meta-analysis reported a 36% increase in risk for each decade decrease in age. Furthermore, a 2025 retrospective cohort study of breast cancer surgery patients corroborated these findings. This association is plausibly attributed to more active neuroplastic and neuroinflammatory processes in younger individuals.58,62 Radiotherapy history and tumor size, in turn, serve as proxies for the extent of tissue and neural damage: adjuvant radiotherapy is among the strongest predictors of pain persisting six months after breast cancer surgery (OR 3.29), while axillary lymph node dissection—compared with sentinel lymph node biopsy—remains the most influential modifiable surgical risk factor (OR 2.41).58,62 Collectively, these findings underscore the need to prioritize preventive analgesic strategies and tailored pain-management protocols for patients whose readily available baseline profile signals high preoperative risk.

Calibration evidence was substantially weaker than discrimination evidence. Although 17 studies reported some calibration information, most relied on plots or goodness-of-fit tests, which do not quantify calibration slope or calibration-in-the-large. Only isolated studies reported mean absolute calibration error, integrated calibration indices, or Brier scores, and the results were too heterogeneous for pooling. Consequently, a moderate AUC should not be interpreted as evidence that predicted absolute risks are accurate.

Comparison with Prior Work

Previous systematic reviews of cancer-related postoperative pain have largely focused on identifying risk factors, evaluating analgesic interventions, or describing prediction models for a single cancer type.63–65 Few have attempted a comprehensive synthesis of all available models across cancer types, and fewer still have conducted quantitative meta-analyses of model performance, with most prior reviews confined to narrative summaries owing to concerns about substantial between-study heterogeneity.58,66,67 Consistent with this literature, we observed marked variation across the included studies in populations, outcome definitions, predictor selection, modeling methodology, and validation strategy. Such fragmentation reflects the absence of standardized research protocols in this field and renders direct head-to-head comparisons of individual models difficult.

Our pooled discrimination metrics align closely with adjacent meta-analytic evidence: a recent systematic review of CPSP prediction models in mixed adult surgical populations reported individual AUROC values from 0.66 to 0.96, a pooled C-index of 0.79, and substantial heterogeneity,67 while the 2024 review of AOPP models found that, although most reported moderate-to-good discrimination, all 17 models were at high risk of bias and only three had undergone external validation.16 The convergence of these reviews with our findings suggests that discriminative performance in this field has plateaued at a clinically suboptimal level, while methodological rigor has failed to keep pace with model proliferation.

Diagnostic performance also differed by phenotype. In validation datasets, pooled sensitivity/specificity were 0.75/0.80 for AOPP and 0.72/0.72 for CPSP. These values indicate only moderate case identification and must be interpreted alongside heterogeneous thresholds and outcome definitions. External-validation evidence was particularly sparse: seven studies reported an external cohort, and none constituted geographically independent validation. AUC, sensitivity, and specificity also do not establish clinical utility; decision-curve analysis and net benefit at actionable thresholds remain necessary before implementation.68

Heterogeneity

Heterogeneity remained substantial in every primary AUC synthesis (total I2 82.3%–97.9%). More importantly, 95% prediction intervals were broad: 0.54–0.95 for training AOPP, 0.62–0.90 for training CPSP, 0.66–0.89 for validation AOPP, and 0.50–0.91 for validation CPSP. These intervals show that an apparently moderate pooled AUC can coexist with poor performance in a new population. Differences in cancer type, surgery, pain definition, prediction window, predictors, and validation method are likely contributors.

Cancer-specific estimates should be interpreted as provisional. Gastrointestinal-cancer AOPP models had pooled AUCs of 0.83 in training and 0.80 in validation and were supported by 10 studies. Breast-cancer CPSP models had lower estimates (0.76 and 0.72). Lung-cancer CPSP models had higher point estimates (0.85 and 0.82), but each synthesis included only three studies and produced wide confidence and prediction intervals. Thus, lung-cancer CPSP appears promising for further validation, not closest to clinical adoption.

Differences across the additional exploratory subgroups should not be interpreted as stable performance advantages because several cells were sparse and the pooled strata often had wide prediction intervals.

The dependence-aware and one-model-per-study analyses produced similar phenotype-specific pooled estimates, and leave-one-study-out analyses did not identify a single dominant study. This robustness does not remove the underlying limitations: most studies were single-center, calibration was incompletely reported, outcome definitions varied, and independent external validation was rare.

Methodological Quality and Risk of Bias

Our PROBAST assessment of the 29 included studies revealed substantial variation in methodological quality, with only 11 of 53 models meeting low-risk criteria and only 7 studies (24.1%) reporting external validation—none in a geographically independent population. These pooled estimates should therefore be interpreted with caution, and most existing models are not yet ready for routine clinical use. Our findings parallel the broader machine-learning prediction-model literature: a systematic review of 152 supervised machine-learning models found 87% (95% CI 81%–91%) at high risk of bias, the analysis domain most often rated high risk, with 56% using inadequate events per candidate predictor, 41% handling missing data inadequately, and 39% assessing overfitting improperly.69 This convergence indicates that the postoperative cancer-pain literature is recapitulating methodological pitfalls seen across other prediction-model domains rather than learning from them.

The principal sources of bias spanned multiple domains. In the participant domain, single-center designs and poor sample representativeness were primary contributors, and failure to report consecutive enrollment further raised the potential for selection bias. In the predictor domain, lack of pre-specification of predictor sets can lead to data dredging and severe overfitting; notably, one study incorporated postoperative pain measurements into a model intended for preoperative risk stratification—a fundamental design flaw that renders the model clinically unusable. In the outcome domain, no study reported blinding of outcome assessors, most relied on retrospective record extraction (susceptible to measurement and recall bias), and several used non-standard CPSP definitions that limit generalizability.

In the analysis domain, insufficient sample size was common, with many studies failing to meet accepted events-per-variable thresholds. The widespread categorization of continuous variables (27 of 29 studies) warrants particular concern: frequently motivated by the wish to facilitate nomogram presentation, it sacrifices statistical information, produces step-functions inconsistent with biological dose–response relationships, and amplifies heterogeneity through arbitrary thresholds. Likewise, the dominance of complete-case analysis (24 of 29 studies) can introduce selection bias when data are not missing completely at random—likely for psychological and pain-related variables. The reliance on EPV ≥ 10 as a sample-size benchmark is itself outdated, having been superseded by data-informed, target-based planning that addresses desired shrinkage, candidate predictor count, functional form, and validation aims. Future studies should therefore adopt tailored, formal sample-size calculations for individual prediction models, as outlined by Riley et al.70

Future priorities differ for AOPP and CPSP. AOPP models should target a clearly defined early postoperative window, use pain thresholds tied to actionable perioperative analgesic decisions, and report sensitivity, specificity, calibration, and decision-curve net benefit. CPSP models require standardized pain definitions and follow-up beyond three months, careful separation of pre-existing pain and recurrence, and prospective multicenter external validation. For both phenotypes, future studies should follow TRIPOD and PROBAST,24 preserve continuous predictors where appropriate, handle missing data rigorously, and evaluate transportability before proposing clinical use.

Limitations

This review has several limitations. Most included studies were single-center and at high or unclear risk of bias. Outcome definitions, prediction windows, and model-development methods varied markedly. Although multilevel models and cluster-robust inference addressed the statistical dependence of multiple models from the same study, they cannot correct design bias or clinical non-comparability. Prediction intervals remained wide, cancer-specific analyses were sparse outside gastrointestinal AOPP and breast CPSP, and bivariate diagnostic analyses included relatively few studies. Calibration could not be pooled because slopes, intercepts, and uncertainty were not reported. Finally, external validation was uncommon and not geographically independent, limiting transportability.

Conclusion

Existing AOPP and CPSP prediction models show moderate average discrimination, but their performance is heterogeneous and uncertain in new settings. No model can currently be recommended for routine clinical use because of wide prediction intervals, high risk of bias, limited calibration evidence, and sparse independent external validation. Gastrointestinal-cancer AOPP models and lung-cancer CPSP models warrant further evaluation, but the latter are supported by only three studies. Future AOPP work should emphasize early actionable risk thresholds, whereas CPSP work should prioritize standardized long-term outcomes and prospective multicenter validation.

Data Sharing Statement

All data analyzed in this systematic review and meta-analysis were derived from previously published, peer-reviewed studies. The datasets supporting the pooled analyses—including the extracted model performance estimates (AUC, sensitivity, and specificity) used to reproduce the meta-analytic findings—are provided in the Multimedia Appendix (Supplementary Tables S3 and S5).

Acknowledgments

The authors thank the experts who provided technical guidance during the preparation of this manuscript.

Author Contributions

Haipeng Xu and Yuqian Li contributed equally as co-first authors and were responsible for study conception, data extraction, statistical analysis, and drafting the manuscript. Xinlei Xu and Haohao Ni performed the literature search, study selection, data extraction, and PROBAST assessment. Yunxing Xie conceived and supervised the study, adjudicated discrepancies, and critically revised the manuscript. All authors made a significant contribution to the work reported, whether that is in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; have agreed on the journal to which the article has been submitted; and agree to be accountable for all aspects of the work.

Funding

This work was supported by the Zhejiang Chinese Medical University 2023 Affiliated Hospital Research Project (grant No. 2023FSYYZQ06) and the Zhejiang Province Traditional Chinese Medicine Science and Technology Project (grant No. 2026ZL0348).

Disclosure

The authors report no conflicts of interest in this work.

References

1. Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229–19. doi:10.3322/caac.21834

2. K PS, Jacob S, Sullivan R, et al. Evidence-based benchmarks for use of cancer surgery in high-income countries: a population-based analysis. Lancet Oncol. 2021;22(2):173–181. doi:10.1016/S1470-2045(20)30589-1

3. M SK, Smith C, Ibadin S, et al. Enhanced recovery after surgery guidelines and hospital length of stay, readmission, complications, and mortality: a meta-analysis of randomized clinical trials. JAMA Network Open. 2024;7(6):e2417310. doi:10.1001/jamanetworkopen.2024.17310

4. F RM, Strang A, Roland G, et al. Perioperative pain management and cancer outcomes: a narrative review. J Pain Res. 2023;16:4181–4189. doi:10.2147/JPR.S432444

5. W BJ, G PA. Influence of opioids on immune function in patients with cancer pain: from bench to bedside. Br J Pharmacol. 2018;175(14):2726–2736. doi:10.1111/bph.13903

6. A SS, P L, Barke A, et al. The IASP classification of chronic pain for ICD-11: chronic postsurgical or posttraumatic pain. Pain. 2019;160(1):45–52. doi:10.1097/j.pain.0000000000001413

7. S KJ, Dana E, X XMZ, et al. Prevalence and risk factors for chronic postsurgical pain after thoracic surgery: a prospective cohort study. J Cardiothorac Vasc Anesth. 2024;38(2):490–498. doi:10.1053/j.jvca.2023.09.042

8. Martinez V, Lehman T, Lavand’homme P, et al. Chronic postsurgical pain: a European survey. Eur J Anaesthesiol. 2024;41(5):351–362. doi:10.1097/EJA.0000000000001974

9. M SBT, Janssen L, C VA, et al. Persistent pain after breast cancer treatment, an underreported burden for breast cancer survivors. Ann Surg Oncol. 2024;31(10):6753–6763. doi:10.1245/s10434-024-15682-2

10. Park R, Mohiuddin M, Arellano R, et al. Prevalence of postoperative pain after hospital discharge: systematic review and meta-analysis. Pain Rep. 2023;8(3):e1075. doi:10.1097/PR9.0000000000001075

11. C RD, M P-ZE. Chronic post-surgical pain - update on incidence, risk factors and preventive treatment options. BJA Educ. 2022;22(5):190–196. doi:10.1016/j.bjae.2021.11.008

12. A SK, D WT, P SS, et al. Predicting chronic postsurgical pain: current evidence and a novel program to develop predictive biomarker signatures. Pain. 2023;164(9):1912–1926. doi:10.1097/j.pain.0000000000002938

13. Xu J, Liu X, Zhao J, et al. Comprehensive review on personalized pain assessment and multimodal interventions for postoperative recovery optimization. J Pain Res. 2025;18:2791–2804. doi:10.2147/JPR.S516249

14. Sun C, Li M, Lan L, et al. Prediction models for chronic postsurgical pain in patients with breast cancer based on machine learning approaches. Front Oncol. 2023;13:1096468. doi:10.3389/fonc.2023.1096468

15. Zhang C, He J, Liang X, et al. Deep learning models for the prediction of acute postoperative pain in PACU for video-assisted thoracoscopic surgery. BMC Med Res Methodol. 2024;24(1):232. doi:10.1186/s12874-024-02357-5

16. Papadomanolakis-Pakis N, V MP, Carlé N, et al. Prognostic clinical prediction models for acute post-surgical pain in adults: a systematic review. Anaesthesia. 2024;79(12):1335–1347. doi:10.1111/anae.16429

17. Papadomanolakis-Pakis N, Uhrbrand P, Haroutounian S, et al. Prognostic prediction models for chronic postsurgical pain in adults: a systematic review. Pain. 2021;162(11):2644–2657. doi:10.1097/j.pain.0000000000002261

18. F WR, M MKG, D RR, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51–58. doi:10.7326/M18-1376

19. Page MJ, Moher D, Bossuyt PM, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160

20. McInnes MDF, Moher D, Thombs BD, et al. Preferred reporting items for a systematic review and meta-analysis of diagnostic test accuracy studies: the PRISMA-DTA statement. Jama. 2018;319(4):388–396. doi:10.1001/jama.2017.19163

21. Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10(1):39. doi:10.1186/s13643-020-01542-z

22. D’Souza RS, Barrington MJ, Sen A, Mascha EJ, Kelley GA. Systematic reviews and meta-analyses in regional anesthesia and pain medicine (Part II): guidelines for performing the systematic review. Reg Anesth Pain Med. 2024;49(6):403–422. doi:10.1136/rapm-2023-104802

23. Barrington MJ, D’Souza RS, Mascha EJ, Narouze S, Kelley GA. Systematic reviews and meta-analyses in regional anesthesia and pain medicine (Part I): guidelines for preparing the review protocol. Reg Anesth Pain Med. 2024;49(6):391–402. doi:10.1136/rapm-2023-104801

24. Moons KG, Altman DG, Reitsma JB, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): explanation and elaboration. Ann Intern Med. 2015;162(1):W1–73. doi:10.7326/M14-0698

25. G MK, De groot J A, Bouwmeester W, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11(10):e1001744. doi:10.1371/journal.pmed.1001744

26. M ŠA. Measures of diagnostic accuracy: basic definitions. Ejifcc. 2009;19(4):203–211.

27. P HJ, G TS, J DJ, et al. Measuring inconsistency in meta-analyses. BMJ. 2003;327(7414):557–560. doi:10.1136/bmj.327.7414.557

28. IntHout J, P IJ, F BG. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014;14:25. doi:10.1186/1471-2288-14-25

29. Egger M, Davey Smith G, Schneider M, et al. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629–634. doi:10.1136/bmj.315.7109.629

30. G AK, M DH, E JH, et al. Predictive factors for the development of persistent pain after breast cancer surgery. Pain. 2015;156(12):2413–2422. doi:10.1097/j.pain.0000000000000298

31. F BL, H ZX, L GB, et al. Predictive model for acute abdominal pain after transarterial chemoembolization for liver cancer. World J Gastroenterol. 2020;26(30):4442–4452. doi:10.3748/wjg.v26.i30.4442

Comments (0)

No login
gif