Agreement in Staging and Treatment Recommendations Among Clinicians, Text-Only Clinicians, DeepSeek-V3, and ChatGPT-4o for Nasopharyngeal Carcinoma Patients.
AI interpretation is pending for this paper.
Open original publication →What the AI sees
Not AI summarized yet.
Research significance
Pending deeper interpretation.
Source abstract
PURPOSE: With AI tools being increasingly utilized for medical inquiries, this study evaluated the agreement between clinicians, text-only clinicians, DeepSeek-V3, and ChatGPT-4o in TNM staging and preferred treatment recommendations for newly diagnosed nasopharyngeal carcinoma (NPC) patients. METHODS: A retrospective study analyzed 322 consecutive NPC patients treated at our institution from January 2023 to February 2025. TNM staging (AJCC 8th edition) and preferred treatment recommendations were independently assessed by text-only clinicians, DeepSeek-V3, and ChatGPT-4o. Interrater agreement was quantified using Cohen's kappa coefficient and Fleiss' kappa analysis (κ), with the range of κ = 0.81-1.00 considered almost perfect agreement. RESULTS: The cohort comprised 244 males and 78 females (median age, 52 years, range, 18-77 years). Cohen's kappa analysis showed that for clinical staging, moderate agreement was exhibited in clinician-AI comparisons, while substantial agreement was observed between clinicians and text-only clinicians. For preferred treatment recommendations, slight agreement was noted in clinician-AI comparisons, while moderate agreement was found between clinicians and text-only clinicians. Fleiss' kappa analysis demonstrated moderate agreement among the 4 raters for T stage, N stage, and clinical staging, while M stage achieved almost perfect agreement. However, overall agreement for treatment recommendations was fair. CONCLUSIONS: The AI tools demonstrated moderate agreement with clinicians in overall clinical staging for NPC, with heterogeneous performance across T, N, and M categories (almost perfect for M stage, moderate for T and N stages), whereas their preferred treatment recommendations showed only slight agreement with clinical decision-making.