Causal estimation and machine learning methods for survival outcomes among AYA cancer patients: A scoping review.
AI interpretation is pending for this paper.
Open original publication →What the AI sees
Not AI summarized yet.
Research significance
Pending deeper interpretation.
Source abstract
Survivors of adolescent and young adult (AYA) cancer experience an elevated risk of premature mortality. While causal inference and machine learning (ML) methods are increasingly applied to observational survival data, the methodological landscape of this work remains unmapped. This scoping review mapped the extent, range, and nature of evidence using these methods to estimate or predict survival outcomes in this population. Following Joanna Briggs Institute (JBI) scoping review methodology, eight databases were searched from inception to November 2025. From 11,584 identified records, 5629 were screened after duplicate removal, of which 104 studies met the inclusion criteria (99 peer-reviewed journal articles and 5 conference abstracts). Of these, 68 (65.4%) were published from 2020 onward, 104 (100.0%) were retrospective cohorts, and most used US registry data (75 [72.1%]; SEER alone, 59 [56.7%]). Eighty-seven studies applied causal inference methods, among which propensity score matching was near-universal (75; 86.2%), followed by inverse probability of treatment weighting (9; 10.3%). None used targeted maximum likelihood estimation, marginal structural models for time-varying confounding, causal forests, or target trial emulation. Methods commonly associated with causal inference were widely used, but formally specified causal analyses were uncommon. Only three studies (3.4%) were designed causal analyses (regression discontinuity, causal mediation, and inverse probability weighted analysis [1 study each]), one (1.1%) provided a complete identification statement, and one (1.1%) reported a sensitivity analysis for unmeasured confounding. Competing risks were relevant where a quantity of specific event(s) was estimated and death from other causes could prevent the observation of specific events. Among the 48 applicable causal studies, 9 (18.8%) used a competing risks or relative survival method. ML prediction methods were used in nineteen studies, most often random survival forests (12; 63.2%); only four (21.1%) were externally validated. Among 99 peer-reviewed studies, the most common methodological issues were competing risks mishandling (22/99; 22.2% overall) and model overfitting (15/99; 15.2%). Research applying causal inference and ML methods to survival outcomes in AYA cancer patients is growing rapidly. However, the evidence draws on a narrow range of methods and data, dominated by propensity score matching using US registry data, with minimal documentation of identifying assumptions, rare probing of unmeasured confounding, and competing risks rarely addressed where the outcome made them relevant. ML models showed adequate discrimination, but a large proportion were not externally validated or calibrated, raising concerns about external validity and generalizability. Closing the gap requires adopting target trial-emulated causal designs with estimand-appropriate competing risk handling, together with externally validated, calibrated mortality prediction tools.