AI-based BCVA prediction across retinal diseases: multimodal mirage or meaningful milestone?
Editorial Commentary

AI-based BCVA prediction across retinal diseases: multimodal mirage or meaningful milestone?

Ezra Eio, Kabilan Elangovan, Daniel Ting

1Singapore National Eye Centre, Singapore Eye Research Institute, Singapore, Singapore; 2Singapore Health Services, Singapore, Singapore; 3Duke-NUS Medical School, Singapore, Singapore

Correspondence to: Daniel Ting, MD (1st Hons), PhD. Associate Professor, Duke-NUS Medical School, Singapore, Singapore; Director, AI Office, Singapore Health Service, Singapore, Singapore; Head, AI and Digital Health, Singapore Eye Research Institute, The Academia, 20 College Road, Level 6 Discovery Tower, Singapore 169856, Singapore. daniel.ting45@gmail.com

Comment on: Dong L, Gao W, Niu L, et al. AI-based prediction of best-corrected visual acuity in patients with multiple retinal diseases using multimodal medical imaging. Br J Ophthalmol 2026;110:158-65.


Keywords: Best-corrected visual acuity (BCVA); prediction; bias; translation; generalizability


Received: 30 May 2026; Accepted: 24 July 2026; Published online: 26 August 2026.

doi: 10.21037/jmai-2026-0095


Introduction

Best-corrected visual acuity (BCVA) remains one of the most fundamental yet imperfect measurements in ophthalmology (1,2). Despite its central role in diagnosis, monitoring, and therapeutic decision-making, BCVA assessment is inherently subjective, dependent on patient response, testing conditions and examiner variability, requiring significant operator time (1,2). These limitations have driven increasing interest in artificial intelligence (AI) approaches that attempt to infer functional vision directly from structural imaging (3,4). While early models have demonstrated promise, most have been restricted to single diseases or single imaging modalities, limiting their relevance to real-world clinical practice (5-7).

In this context, the study by Dong et al. represents a meaningful methodological advance (8). By integrating macular optical coherence tomography (OCT), optic disc OCT, and fundus imaging across a heterogeneous retinal cohort, the authors move toward a more realistic representation of clinical complexity, with multimodal and multi-disease incorporation. Their reported internal performance is striking, achieving a mean absolute error (MAE) of 2.865 and R2 of 0.935, as reported in the original study. MAE measures the average magnitude of errors between the actual and predicted BCVA values, serving as an indicator of how accurate or how far off model predictions are. It is calculated per eye and expressed in ETDRS letters. However, in external prospective validation, MAE increases to 8.38—revealing a clear disparity in performance (8). The corresponding external R2 was not reported, an omission that itself warrants critical discussion, since it prevents assessment of how much explanatory power was lost on external data.

This gap reflects a wider challenge in AI in medicine, where strong algorithmic and model performance may not necessarily translate into real-world clinical impact. We argue that the study constitutes a genuine methodological advance in multimodal BCVA prediction, although steps must be taken to improve its reliability for routine clinical translation. Despite the considerable potential of AI-based multimodal prediction of BCVA, several important limitations—including dataset bias (data collection), uncertain generalizability (model development), limited alignment, translation into clinical workflows and equitable deployment (clinical integration) must first be addressed. Taken together, these barriers emphasise the need to move beyond algorithmic, accuracy-focused metrics towards holistic evaluation frameworks which encompass clinical utility, safety and responsible implementation.


Dataset bias and the illusion of performance in heterogeneous AI models

A defining strength of the study is its inclusion of multiple retinal diseases, moving beyond narrow, single-disease modelling (8). However, this heterogeneity is undermined by substantial imbalance in dataset composition. Refractive error (RE) accounts for approximately 70% of the dataset (731 of 1,002 of cases) (Figure 1) (8) representing cases that typically exhibit near-normal visual acuity and minimal structural pathology. These are inherently easier to predict, and as a result, aggregate metrics such as MAE are disproportionately influenced by low-complexity cases. RE shows the highest median visual acuity (85 ETDRS letters vs. 58–65 for other diseases) (Figure 2).

Figure 1 Dataset distribution. Dataset composition and model performance across six ophthalmic diseases in an AI-based BCVA prediction study. (A) Distribution of eyes across training, test, internal validation and external prospective datasets. (B) BCVA distribution in the external prospective cohort. (C) Model performance stratified by disease based on the best-performing three-modality AA model on the internal test set. AA, average-aggregation; AI, artificial intelligence; BCVA, best-corrected visual acuity.
Figure 2 Development gap in AI-based BCVA prediction across imaging modality configurations and validation cohorts. (A) MAE stratified by test set, internal validation set, and external prospective cohort. (B) R2 values for internal datasets only; external R2 was not reported. AI, artificial intelligence; BCVA, best-corrected visual acuity; MAE, mean absolute error; OCT, optical coherence tomography; OD, optic disc.

Disease-stratified analysis reveals a markedly different narrative. While RE achieves near-ceiling performance (MAE 1.063), prediction errors are significantly higher for clinically meaningful conditions such as age-related macular degeneration (AMD; MAE 9.493), diabetic retinopathy (DR; MAE 10.111), and retinal vein occlusion (RVO; MAE 8.262) (Figure 1) (8). This discrepancy is critical. The clinical value of BCVA prediction lies not in confirming normal vision, but in accurately characterising impaired vision where treatment decisions depend on subtle but meaningful differences (9-11). To illustrate how the dataset imbalance inflates the aggregate MAE figure, an approximate weighted average using the reported RE proportion (73%) and the mean MAE of the three non-RE diseases above (27%, mean MAE =9.29) yields an estimated MAE of 3.3. This is close to the published aggregate of 2.865 and confirms that the aggregate figure is driven by the RE majority rather than by performance on visually significant disease.

This issue is further exacerbated by the underrepresentation of severe visual impairment. Only 2.63% of the dataset falls within low BCVA ranges (ETDRS score less than 30) (8), despite this being the subgroup where prediction accuracy is most consequential. The model's reduced performance in this group suggests insufficient representation during training, reinforcing a broader pattern in medical AI: models often perform best where prediction is easiest, rather than where it is most needed (10).

While big data in ophthalmology shows promise in transforming clinical research, how the study population is defined must be closely examined and as representative of the general population as possible. The systemic skew of AI models due to non-representative training data remains a major barrier. Methods such as propensity score matching, inverse probability weighting and stratifications help to reduce selection bias (12).

A deeper concern relates to how the aggregate R2 score of 0.935 is interpreted. This figure is substantially inflated by the dense cluster of RE cases with near-normal BCVA values, which compress variance around the high end of the scale. Given the large proportion of RE cases within the dataset, the resulting R2 figure will be assessed based on a narrower, unrepresentative subgroup with narrower variance. This produces optimistic estimates that do not generalize to the target clinical population (3,4,13). The authors acknowledge poor low-BCVA prediction, yet present headline metrics without adequately qualifying this limitation, a framing that risks misleading readers unfamiliar with the compositional asymmetries in the underlying data (8).

From a data ethics and transparency perspective, this raises important questions. Aggregate reporting obscures performance disparities across clinically relevant subgroups, potentially misleading end-users regarding model reliability. Reporting frameworks such as TRIPOD-AI advocate for subgroup-stratified performance disclosure as a minimum standard for clinical AI publications (3,4,14). Ethical translation of such models will also require documented patient consent for the secondary use of imaging data in model development, pre-specified evaluation of algorithmic bias across age, sex and ethnicity, and a defined regulatory pathway prior to clinical deployment. A shift toward mandatory disease-stratified and severity-stratified reporting is therefore essential. Without such standards, clinicians may be forced to rely on aggregate performance metrics that inadequately reflect real-world clinical utility, particularly in the patient populations where these systems are most likely to be deployed.


Clinical misalignment: predictive accuracy may not translate to clinical utility

The clinical validity of the prediction model requires further examination. In this study, BCVA values are derived from decimal visual acuity and converted into ETDRS scores using a mathematical transformation (2,8). While practical, this introduces systematic noise into the ground truth. Variability in visual acuity measurement is well documented (1,2), and non-standardised testing conditions further compound this issue. As a result, the model is trained to reproduce an inherently noisy signal, placing an upper bound on achievable performance.

More importantly, the interpretation of prediction error is not contextualised within clinical decision-making thresholds. The study reports that over 92% of predictions fall within 10 ETDRS letters of the ground truth. It is not clear whether this figure reflects all diseases equally or is disproportionately driven by the RE-dominated cohort, since the original study does not report disease-stratified accuracy, which is a further reporting gap. However, a 10-letter difference corresponds to a clinically meaningful change, a threshold widely used in therapeutic trials such as RISE, RIDE, VISTA, and VIVID (9,13,15). While BCVA is rarely the sole criterion for treatment decisions, as structural findings, patient-reported outcomes, and clinician judgement all contribute, even modest inaccuracies of this magnitude could contribute to inappropriate management when combined with other clinical factors. Reporting that 92% of predictions fall within this margin therefore implies that up to 8% may not, a rate that warrants caution given how often BCVA continues to inform treatment thresholds in practice (13).

This distinction between statistical performance and minimum clinically important difference (MCID) is a recurring blind spot in medical AI evaluation. A model that reliably predicts BCVA as 68 letters when the true value is 72 may appear accurate by MAE standards yet would misclassify a patient who has crossed the threshold for treatment initiation. The study does not report sensitivity or specificity at any clinical action threshold, nor does it assess the model’s ability to correctly identify treatment-eligible versus non-eligible eyes (9,15). This represents a fundamental gap between what the model measures and what clinicians need it to answer.

The absence of uncertainty quantification further limits clinical applicability. The model provides point estimates without confidence intervals, offering no indication of prediction reliability on a case-by-case basis. In clinical settings, this is a critical limitation. To ensure safe deployment of medical AI, safeguards for cases involving rare diseases, suboptimal image quality or atypical presentations is needed, such as flagging up if the model is uncertain (4). Without calibrated confidence outputs, the model cannot flag cases for human review, undermining its value as a safe adjunct tool.

A common counterargument is that such models are intended as screening or triage tools rather than replacements for clinical assessment. While valid, even adjunct tools must be evaluated against clinically meaningful benchmarks. The absence of prospective head-to-head comparison with standard clinical BCVA testing leaves the model’s incremental utility entirely unquantified, making it impossible to assess whether AI-derived BCVA estimates would improve, replicate, or degrade clinical decision-making in practice (3,5).


Multimodal AI and deployment equitability in healthcare systems

The study demonstrates that multimodal fusion improves predictive performance, reflecting the complementary nature of structural imaging modalities (8,16). OCT provides high-resolution cross-sectional detail of retinal microstructure (16), while fundus imaging offers broader contextual information including vascular architecture and macular integrity. Together, they enable a more comprehensive representation of retinal pathology (17,18).

However, this technical advantage introduces a practical paradox. The incremental performance gain achieved by adding the third modality (optic disc OCT) over the two-modality model is approximately one ETDRS letter (MAE 3.86 versus 2.865) (8). This marginal gain must be weighed against the operational cost of acquiring an additional imaging modality in every patient. Clinically, the question is not whether three modalities outperform two in a controlled experiment, but whether that one-letter improvement justifies the added workflow burden, imaging time, and infrastructure investment at scale. Formal cost-effective analysis, for example, estimating the incremental cost per additional imaging modality against the diagnostic benefit gained, was not undertaken in the original study. This is an important gap for health-eocnomic evaluation and is a central question for any health technology assessment (10). Cost-benefit analysis of multimodal imaging with analytical frameworks (e.g., MCID for AI models) (19) provides key information to guide the usage of multimodal BCVA prediction.

From a healthcare integration perspective, the infrastructure requirements of this model are substantial. Both the Heidelberg Spectralis SD-OCT and Topcon non-mydriatic fundus camera used in this study are high-cost, fixed-installation devices found predominantly in tertiary referral centres (8). The study's suggestion that the model offers utility in resource-limited settings sits in direct tension with this reality (8,10). Globally, the populations with the greatest unmet need for retinal disease detection, including those in low- and middle-income countries where AMD, DR, and RVO are rapidly increasing in prevalence (10,11), are precisely those least likely to have access to such infrastructure. A model designed around spectral-domain OCT cannot, by definition, serve these populations without fundamental redesign.

This reflects a broader structural issue in medical AI development whereby the development and optimisation environment systematically differs from the deployment environment. Models trained and validated in well-resourced academic centres embed assumptions about imaging quality, device consistency, and clinical workflow that simply do not transfer to district hospitals or community screening programmes. Approaches such as federated learning, domain adaptation, and device-agnostic model architectures have been proposed as partial solutions (3,8), but none are evaluated here. Tiered modelling strategies, in which lightweight single-modality models are explicitly designed, validated, and benchmarked for low-resource settings rather than treated as fallback options, represent a more principled path toward equitable deployment (4,10).


Generalizability and the reproducibility challenge in medical AI

The most important insight from the study emerges from its external validation results. While internal testing demonstrates strong performance (MAE 2.865), MAE increases to 8.38 in the external prospective cohort (8), reflecting substantial performance degradation primarily because the external dataset contains a higher proportion of low-BCVA cases (36.35% with ETDRS score below 60, versus 11.83% in the internal test set) (Figure 2). It also remains unclear whether this external cohort was drawn from an independent institution and population, or instead represents a temporally distinct cohort from the same centre. The sample size and recruitment procedure for this cohort are not fully described in the primary report, which limits assessment of whether it adequately represents the intended real-world use population beyond case mix differences alone.

The authors attribute this degradation to case-mix differences, and this explanation is plausible. However, it also reveals a deeper problem: the model has not learned to predict visual acuity from retinal structure in a general sense, but rather to predict the BCVA distribution of a specific training cohort. When that distribution shifts, predictive capacity deteriorates sharply (3,8). This manifestation of covariate shift, reveals the imbalanced and inadequate nature of the datasets, and reveals the poor accuracy and generalizability of the model in predicting BCVA (3,4).

The rise of big data in ophthalmology with its large cohort size and varied participant characteristics (3 “V”s: volume, variety, velocity) (20) is a potential strategy to improve generalizability. An example is the UK Biobank, a longitudinal cohort study of 500,000 participants, with extensive oculomic data which may improve model training and internal and external validation.

A further underappreciated limitation is that the external validation dataset, while prospective, remains device-homogeneous. Both training and external validation used the same Heidelberg Spectralis OCT and Topcon fundus camera platforms (8). Real-world deployment would inevitably involve Zeiss, Optovue, Canon, and other device ecosystems, each with distinct image characteristics, layer segmentation protocols, and signal-to-noise profiles. Recent work has directly demonstrated reduced deep learning segmentation performance when models trained on Heidelberg Spectralis images are applied to Zeiss Cirrus scans without device-specific adaptation (21), underscoring that this is not a theoretical concern. The true extent of device-induced performance degradation remains unquantified by this study. Given that image acquisition variability is among the most potent sources of domain shift in ophthalmic AI (3,4,22), this is a significant omission that substantially limits confidence in the model’s readiness for broad clinical deployment.

To enhance reproducibility, code availability and sharing has been increasingly adopted and practised (8). Beyond code availability, the datasets on which models are built must be subject to robust and standardised data collection protocols, with adequate disease diversity and population representation. The opacity of training data, which remains available only on request and subject to data-sharing agreements, further limits independent verification (3,4,8). These are systemic challenges for the field, and addressing them will require coordinated infrastructure investment rather than individual study-level effort alone.

Emerging AI paradigms may help address some limitations. Foundation models pretrained on large unlabelled retinal images using self-supervised learning such as RETFound have shown improved label efficiency and cross-domain generalizability compared to supervised models. Synthetic data generation may help rebalance datasets with under-represented disease types and severity, without relying on proportionally larger real-world datasets. These emerging AI models merit consideration in further iterations of multimodal BCVA prediction models.


Conclusions

The study by Dong et al. is an important step forward for multimodal AI in BCVA prediction, demonstrating that structural retinal imaging can be an approximation for visual function in controlled circumstances. However, it also highlights the major challenges that must be addressed before such models can be translated into clinical practice.

Dataset imbalances hide clinically relevant performance, while the noisy ground truth and misaligned evaluation metrics hinder interpretability. Multimodal designs, can provide technical advantages, but there are practical limitations that prevent their real-world deployment. Most importantly, the significant drop in external validation performance highlights the fragility of current models when exposed to clinical heterogeneity.

Future work should focus on quality, transparency, and diversity of data, alongside robust multi-centre validation. To address dataset bias and enhance clinical plausibility of the results, future studies can report MAE stratified by disease severity i.e., RE vs. DR, AMD; and severity quartile i.e., mild-moderate vs. severe visual impairment. Evaluation frameworks should move beyond aggregate accuracy metrics to include measures that are clinically meaningful. Incorporating uncertainty estimation using techniques such as Monte Carlo dropout, and aligning models with real world resource constraints will be essential for safe and effective deployment.

Ultimately, the salient question is not whether AI systems can approximate BCVA, but whether they can do so with sufficient reliability and reproducibility to have a meaningful impact on clinical decision-making across a range of healthcare settings. Overcoming this challenge will require more than incremental improvements in model performance; it will require a broader shift in the definition of success in medical AI, moving beyond numerical accuracy to demonstrate clinical value and real-world applicability. Until these challenges are addressed, multimodal BCVA prediction should be regarded as a promising yet still transitional step toward the meaningful integration of AI into ophthalmic practice.

This manuscript has not been previously presented or published in any form. Claude Code was used to generate the figures.


Footnote

Provenance and Peer Review: This article was commissioned by the editorial office, Journal of Medical Artificial Intelligence. The article has undergone external peer review.

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0095/prf

Funding: None.

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0095/coif). D.T. was supported by the NMRC, Duke-NUS Medical School and ASTAR; they are not directly related to this manuscript. The other authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Ricci F, Cedrone C, Cerulli L. Standardized measurement of visual acuity. Ophthalmic Epidemiol 1998;5:41-53. [Crossref] [PubMed]
  2. Baker CW, Josic K, Maguire MG, et al. Comparison of Snellen Visual Acuity Measurements in Retinal Clinical Practice to Electronic ETDRS Protocol Visual Acuity Assessment. Ophthalmology 2023;130:533-41. [Crossref] [PubMed]
  3. Gulshan V, Peng L, Coram M, et al. Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs. JAMA 2016;316:2402-10. [Crossref] [PubMed]
  4. Ting DSW, Liu Y, Burlina P, et al. AI for medical imaging goes deep. Nat Med 2018;24:539-40. [Crossref] [PubMed]
  5. Paul W, Burlina P, Mocharla R, et al. Accuracy of Artificial Intelligence in Estimating Best-Corrected Visual Acuity From Fundus Photographs in Eyes With Diabetic Macular Edema. JAMA Ophthalmol 2023;141:677-85. [Crossref] [PubMed]
  6. Ye X, Gao K, He S, et al. Artificial Intelligence-Based Quantification of Central Macular Fluid Volume and VA Prediction for Diabetic Macular Edema Using OCT Images. Ophthalmol Ther 2023;12:2441-52. [Crossref] [PubMed]
  7. Wei L, He W, Wang J, et al. An Optical Coherence Tomography-Based Deep Learning Algorithm for Visual Acuity Prediction of Highly Myopic Eyes After Cataract Surgery. Front Cell Dev Biol 2021;9:652848. [Crossref] [PubMed]
  8. Dong L, Gao W, Niu L, et al. AI-based prediction of best-corrected visual acuity in patients with multiple retinal diseases using multimodal medical imaging. Br J Ophthalmol 2026;110:158-65. [Crossref] [PubMed]
  9. Aiello LP, Beck RW, Bressler NM, et al. Rationale for the diabetic retinopathy clinical research network treatment protocol for center-involved diabetic macular edema. Ophthalmology 2011;118:e5-14. [Crossref] [PubMed]
  10. GBD 2019 Blindness and Vision Impairment Collaborators, Vision Loss Expert Group of the Global Burden of Disease Study. Causes of blindness and vision impairment in 2020 and trends over 30 years, and prevalence of avoidable blindness in relation to VISION 2020: the Right to Sight: an analysis for the Global Burden of Disease Study. Lancet Glob Health 2021;9:e144-60. [Crossref] [PubMed]
  11. Guymer RH, Campbell TG. Age-related macular degeneration. Lancet 2023;401:1459-72. [Crossref] [PubMed]
  12. Lee CS. Uses and Abuses of Big Data in Ophthalmology Research. Ophthalmology 2026;133:825-8. [Crossref] [PubMed]
  13. Baxter SL, Kim JE. Artificial Intelligence for Visual Acuity-Gaps From Algorithm to Actualization. JAMA Ophthalmol 2023;141:685-6. [Crossref] [PubMed]
  14. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
  15. Nguyen QD, Brown DM, Marcus DM, et al. Ranibizumab for diabetic macular edema: results from 2 phase III randomized trials: RISE and RIDE. Ophthalmology 2012;119:789-801. [Crossref] [PubMed]
  16. Aumann S, Donner S, Fischer J, et al. Optical Coherence Tomography (OCT): Principle and Technical Realization. 2019; [PubMed]
  17. Liu TYA, Ling C, Hahn L, et al. Prediction of visual impairment in retinitis pigmentosa using deep learning and multimodal fundus images. Br J Ophthalmol 2023;107:1484-9. [Crossref] [PubMed]
  18. Shi XH, Ju L, Dong L, et al. Deep Learning Models for the Screening of Cognitive Impairment Using Multimodal Fundus Images. Ophthalmol Retina 2024;8:666-77. [Crossref] [PubMed]
  19. Draak THP, de Greef BTA, Faber CG, et al. The minimum clinically important difference: which direction to take. Eur J Neurol 2019;26:850-5. [Crossref] [PubMed]
  20. Soh ZD, Cheng CY. Application of big data in ophthalmology. Taiwan J Ophthalmol 2023;13:123-32. [Crossref] [PubMed]
  21. Mukherjee S, De Silva T, Duic C, et al. Validation of Deep Learning-Based Automatic Retinal Layer Segmentation Algorithms for Age-Related Macular Degeneration with 2 Spectral-Domain OCT Devices. Ophthalmol Sci 2025;5:100670. [Crossref] [PubMed]
  22. Varadarajan AV, Bavishi P, Ruamviboonsuk P, et al. Predicting optical coherence tomography-derived diabetic macular edema grades from fundus photographs using deep learning. Nat Commun 2020;11:130. [Crossref] [PubMed]
doi: 10.21037/jmai-2026-0095
Cite this article as: Eio E, Elangovan K, Ting D. AI-based BCVA prediction across retinal diseases: multimodal mirage or meaningful milestone? J Med Artif Intell 2026;09:73.

Download Citation