Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.
This systematic review and meta-analysis of 53 studies reports proof-of-concept AI performance for neuroblastoma diagnosis, risk stratification, prognosis, and genomic characterization, while finding inadequate chemotherapy-response prediction and major deficits in calibration and external validation.
Open original publication →What the AI sees
This systematic review and meta-analysis of 53 studies reports proof-of-concept AI performance for neuroblastoma diagnosis, risk stratification, prognosis, and genomic characterization, while finding inadequate chemotherapy-response prediction and major deficits in calibration and external validation.
Research significance
Evidence in the review suggests that hybrid nomograms, AI-derived prognostic nomograms, and gene signatures may improve neuroblastoma risk or prognosis estimation; inferentially, prospectively validated models could support treatment selection or surveillance, but the supplied record does not demonstrate improved treatment outcomes or clinical utility.
Source abstract
PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.