Interpretability and external validation remain key barriers to clinical integration of multimodal deep learning for breast cancer prognosis
The review by Maigari et al. provides a clear and well-organized look at the current direction of multimodal deep learning for breast cancer prognosis (1). As radiologists, we are seeing firsthand how these tools are moving from theoretical frameworks to early-stage clinical exploration, and several points from the authors’ discussion resonate with challenges faced in daily practice—particularly in middle-income healthcare systems.
One of the recurring issues is the dependence on large public datasets such as The Cancer Genome Atlas (TCGA) and the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC). Although these datasets have helped move the field forward, the reality in most radiology departments is quite different. Scanner models vary, protocols change from site to site, and genomic or pathology data are not always available. Because of this, a model that performs well on METABRIC or TCGA may not behave the same way elsewhere (2). Several studies have already shown that even strong artificial intelligence (AI) systems can lose accuracy when tested in patient populations that differ from the ones used during model development.
Another central theme in the review is the strong performance of multimodal convolutional neural network (CNN)-based and ensemble approaches. While these results are promising, most of these architectures still offer very limited interpretability. For clinicians, especially those navigating multidisciplinary tumor boards or high-volume reading rooms, understanding why a model assigns a particular risk category is just as important as the numerical output itself. Broader analyses in medical AI continue to emphasize that adoption depends not only on accuracy but on clarity, reliability, and the ability of the system to justify its predictions in clinically meaningful terms (3).
The authors also mention the scarcity of external validation across independent cohorts. This remains one of the largest barriers to clinical integration. Testing these models across diverse scanners, patient populations, and healthcare systems—especially those outside high-resource settings—would provide a more realistic understanding of their robustness. Expanding collaborative validation initiatives would be an important next step.
Overall, the review by Maigari et al. gives a clear picture of where multimodal deep learning currently stands in breast cancer prognosis, and it also shows how quickly the field is moving. What seems most important going forward is not only improving model performance, but making these systems easier to understand and more reliable across different clinical settings. Better transparency, more attention to day-to-day data variability, and stronger external validation will likely determine whether these models can move from research projects into routine oncology practice.
Acknowledgments
The authors would like to state that no additional contributors were involved in this work. This manuscript has not been previously presented at any meeting or conference.
Footnote
Provenance and Peer Review: This article was a standard submission to the journal. The article did not undergo external peer review.
Funding: None.
Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-1-258/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Maigari A. XinYing C, Zainol Z. Multimodal deep learning breast cancer prognosis models: narrative review on multimodal architectures and concatenation approaches. J Med Artif Intell 2025;8:61.
- Martin-Noguerol T, Luna A. External validation of AI algorithms in breast radiology: the last healthcare security checkpoint? Quant Imaging Med Surg 2021;11:2888-92. [Crossref] [PubMed]
- Rajpurkar P, Chen E, Banerjee O, et al. AI in health and medicine. Nat Med 2022;28:31-8. [Crossref] [PubMed]
Cite this article as: Bermúdez IAB, Rivera HA. Interpretability and external validation remain key barriers to clinical integration of multimodal deep learning for breast cancer prognosis. J Med Artif Intell 2026;9:40.

