Critical appraisal of standalone AI versus radiologist interpretation in unilateral surveillance mammography after mastectomy
Editorial Commentary

Critical appraisal of standalone AI versus radiologist interpretation in unilateral surveillance mammography after mastectomy

Muharrem Oner ORCID logo, Kefah Mokbel

London Breast Institute, The Princess Grace Hospital, London, UK

Correspondence to: Muharrem Oner, MD. Breast Surgery Unit, London Breast Institute, The Princess Grace Hospital, Barker Gillette Llp, Harley St, London W1G 9QP, UK. Email: muharremoner@gmail.com.

Comment on: Ha SM, Lee JM, Jang MJ, et al. Breast Cancer Detection with Standalone AI versus Radiologist Interpretation of Unilateral Surveillance Mammography after Mastectomy. Radiology 2025;315:e242955.


Keywords: Breast cancer; artificial intelligence (AI); surveillance; mammography


Received: 16 April 2025; Accepted: 11 July 2025; Published online: 23 August 2025.

doi: 10.21037/jmai-2025-110


The recently published article by Ha and colleagues compared standalone artificial intelligence (AI) with radiologist interpretation for breast cancer detection in unilateral surveillance mammography after mastectomy (1). The authors reported that standalone AI demonstrated higher sensitivity but lower specificity. While the study provides valuable insights into the potential role of AI in this specific clinical context, several methodological limitations and analytical concerns deserve further consideration.

First, the retrospective single-institution design introduces selection bias and limits generalizability. The authors excluded 834 of 5,018 eligible patients (16.6%), including 519 with prior contralateral breast surgery. This exclusion is problematic as it creates an artificial study population that does not reflect real-world clinical practice where postsurgical changes are common among patients with a personal history of breast cancer.

Second, there appears to be a fundamental mismatch in the comparison methodology. The radiologists had access to prior examinations for comparison, while the AI system did not. This differential advantage for radiologists may have actually underestimated AI’s comparative performance, yet paradoxically, the authors still found higher sensitivity for AI. This raises questions about whether radiologists were underperforming in this cohort or if other factors influenced interpretation.

Third, the threshold selection for AI positivity (scores ≥10) merits further scrutiny. The authors do not sufficiently justify this threshold or explore how different thresholds might affect the sensitivity-specificity balance. Without a proper receiver operating characteristic (ROC) analysis across different thresholds, it is difficult to assess the optimal operating point for this specific application.

Fourth, the authors report a remarkably low recall rate for radiologists (3.3%) compared to typical rates in surveillance settings, which suggests either exceptional radiologist performance or unique institutional practices that may not translate elsewhere. This unusual finding warrants a more extensive explanation.

Fifth, while AI detected 16 cancers missed by radiologists, it did so at the cost of 193 additional false positives. The clinical implications of this trade-off are inadequately addressed. The positive predictive value of AI (17.3%) was significantly lower than that of radiologists (44.2%), which has substantial workflow and resource implications that are not thoroughly discussed.

Finally, the critical finding that 30.6% of cancers were missed by both radiologists and AI deserves greater emphasis. This highlights that current AI technology, at least as implemented in this study, is far from a comprehensive solution for surveillance challenges. However, in an independent study, an AI system for breast cancer screening outperformed six radiologists, showing an 11.5% higher area under the curve (AUC)-ROC than the average human reader. When simulated as part of the UK double-reading workflow, the AI maintained non-inferior accuracy while reducing the second reader’s workload by 88%, supporting its potential to enhance screening efficiency and accuracy in clinical trials (2).

These limitations do not negate the value of this research but suggest caution in interpreting the results and their clinical applicability.

An important avenue for future research that was not addressed is the development and evaluation of multimodality AI systems. The authors noted that many cancers missed by both radiologists and AI were subsequently detected by ultrasound or magnetic resonance imaging (MRI). This highlights the potential benefit of AI systems that can integrate and analyze data across multiple imaging modalities simultaneously, as demonstrated in recent studies by Yala et al. (3) and Aboutalib et al. (4). Multimodality AI could potentially address the 30.6% of cancers missed by both mammography interpretation methods, particularly in patients with dense breasts. Such integrated approaches might better mimic the clinical reality where radiologists often synthesize information from multiple imaging studies, potentially achieving the improved cancer detection rates shown by Marinovich et al. (5) in their systematic review of combined imaging modalities.

Future studies should consider prospective designs, multi-institutional settings, inclusion of patients with postsurgical changes, matched comparison conditions for AI and radiologists, threshold optimization analyses, multimodality AI integration, and comprehensive assessments of the cost-benefit balance of implementing AI in surveillance workflows.


Acknowledgments

None.


Footnote

Provenance and Peer Review: This article was a standard submission to the journal. The article has undergone external peer review.

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-110/prf

Funding: None.

Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-110/coif). K.M. is a fractional shareholder of Datar Genetics’ stock. K.M. has received honoraria for offering academic and clinical advice to Merit Medical and QMedical corporations. Furthermore, he owns shares in HCA Healthcare UK. The other author has no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Ha SM, Lee JM, Jang MJ, et al. Breast Cancer Detection with Standalone AI versus Radiologist Interpretation of Unilateral Surveillance Mammography after Mastectomy. Radiology 2025;315:e242955. [Crossref] [PubMed]
  2. McKinney SM, Sieniek M, Godbole V, et al. International evaluation of an AI system for breast cancer screening. Nature 2020;577:89-94. [Crossref] [PubMed]
  3. Yala A, Mikhael PG, Strand F, et al. Toward robust mammography-based models for breast cancer risk. Sci Transl Med 2021;13:eaba4373. [Crossref] [PubMed]
  4. Aboutalib SS, Mohamed AA, Berg WA, et al. Deep Learning to Distinguish Recalled but Benign Mammography Images in Breast Cancer Screening. Clin Cancer Res 2018;24:5902-9. [Crossref] [PubMed]
  5. Marinovich ML, Hunter KE, Macaskill P, et al. Breast Cancer Screening Using Tomosynthesis or Mammography: A Meta-analysis of Cancer Detection and Recall. J Natl Cancer Inst 2018;110:942-9. [Crossref] [PubMed]
doi: 10.21037/jmai-2025-110
Cite this article as: Oner M, Mokbel K. Critical appraisal of standalone AI versus radiologist interpretation in unilateral surveillance mammography after mastectomy. J Med Artif Intell 2026;9:10.

Download Citation