How far can DeepSeek-R1 and OpenAI see in ophthalmology diagnosis and decision making?
The manuscript published by Mikhail et al., compared two large language models (LLMs), DeepSeek-R1 and OpenAI o1, for their diagnostic accuracy and decision making across ophthalmology subspecialties (1). LLMs have rapidly and widely been adopted in medicine for clinical and educational purposes (2-4). With the growing interest in LLMs, Mikhail et al. brings us closer to answering a fundamental question: can LLMs earn a seat at the slit-lamp, or do they pose a risk for exacerbating medical errors in ophthalmology?
Behind the lens: a quick look at LLMs in medicine
LLMs are text-based generative artificial intelligence (AI) tools capable of performing natural language processing tasks, often without needing to be trained on specific clinical scenarios (5).They are commonly divided into two broad categories based on their accessibility: proprietary models are closed and company-controlled, whereas open-weight models are publicly available, downloadable, and modifiable (3). Currently, proprietary models dominate the field, with 72% of surveyed LLM users relying on models like OpenAI (6). As the field evolves rapidly, the evaluation of new open-weight models like DeepSeek-R1 is becoming increasingly critical. Mikhail et al.’s study therefore provides timely insights into the diagnostic accuracy, management decision-making, and associated costs of DeepSeek-R1.
Clinical reasoning is a complex cognitive process developed over years of training and pattern recognition by physicians, making one wonder whether such intuition and decision-making skills could ever be fully replicated by AI (7,8). Previous studies have provided promising results on clinical data synthesis and diagnostic performance with proprietary models from companies like OpenAI, Google, and Anthropic (9-13). Ophthalmology has been one of the top five specialties studied for LLM applications (10,14,15). Since margin for error in medicine is very low, continuous monitoring and evaluation are essential to ensure LLM reliability and optimize patient safety and outcomes (16).
As promising as LLMs are, they rely heavily on extensive training data, bringing data breaches, privacy risks, and regulatory compliance challenges to the forefront (17). Medicine inherently involves sensitive protected health information (PHI) governed by privacy regulations like Health Insurance Portability and Accountability Act (HIPAA), and these challenges are further magnified in fields such as ophthalmology (18). For example, fundus and iris photographs contain biometric identifiers that could compromise patient confidentiality if mishandled by external AI services. While the U.S. Food and Drug administration (FDA) is actively working on a dedicated regulatory framework for AI-based technologies, currently LLMs used in medical devices and applications primarily remain under medical device regulations (19).
Welcoming DeepSeek-R1
DeepSeek-R1, a reasoning-first open-weight LLM model released in January 2025, represents a significant leap forward in the capabilities of open-weight models (20,21). Before its launch, open-weight models lagged behind proprietary systems in accuracy and hallucination control (22,23). This was largely due to their reliance on public training datasets, limited computational resources, and the absence of controlled environments for model optimization (24).
DeepSeek-R1 has narrowed this gap, trailing the top-performing proprietary model (o3-mini) by only 2 % on the top-level MATH benchmark (22). This success is attributed to its multistage reinforcement learning framework integrated with Chain-of-Thought generation, which can mitigate the occurrence of hallucinations (25). Moreover, with fine-tuning, DeepSeek-R1 can be specifically adapted for ophthalmology applications. Fine-tuning, already a proven approach for proprietary models, has demonstrated success in improving the accuracy of generated ophthalmology operation notes, as noted in one study on ChatGPT (26).
A systematic review revealed that, before 2025, only 13.4% of LLM studies in ophthalmology analyzed open-weight models (27). Post-launch investigations have demonstrated that DeepSeek-R1 can analyze patient data, provide differential diagnoses and offer guideline-based treatment recommendations (28). As a highly adaptable open-weight model, DeepSeek-R1 also allows institutions to deploy locally (20,29). This ensures that patient data remains within protected internal healthcare systems, addressing critical data privacy concerns by ensuring that PHI remains within the covered healthcare entity (20,29).
DeepDive into the work of Mikhail et al.
DeepSeek-R1 is rapidly gaining recognition, and Mikhail et al. wisely acted promptly to evaluate its utility within ophthalmology, addressing key gaps in the literature. For instance, prior studies have disproportionately concentrated on comprehensive ophthalmology, leaving certain subspecialties relatively underexplored (27,30). This could be the reason why OpenAI models excel in answering general ophthalmology exam questions but demonstrate reduced effectiveness in subspecialty domains (11,31,32). Additionally, fewer than 5% of studies have evaluated the clinical decision-making capacity of LLMs, further highlighting a critical gap in LLM applications (27).
To address these gaps in the literature, Mikhail et al. analyzed 422 complex cases of ten subspecialties, including several underrepresented domains such as neuro-ophthalmology (16%), anterior segment conditions (cornea, cataract, and lens disorders) (11%), uveitis (16%), pathology and tumors (10%), and oculoplastic interventions (7%). Retina and vitreous cases accounted for the largest share at 23%, reflecting its prominence in earlier literature in the field (27). By intentionally incorporating nearly 50% of their data from traditionally overlooked subspecialties, Mikhail et al.’s study makes strides toward a more comprehensive and realistic representation of real-world ophthalmology practice. However, the relatively small sample size underscores the need for larger datasets encompassing the full spectrum of disease presentations and stages before clinical deployment can be considered.
The cases were presented in text format and included at least one image description, capturing visual elements critical for diagnosis. The authors employed Zero-Shot, Plan-and-Solve Plus (PS+) prompting strategy, which breaks tasks into sequential subtasks and refines outputs (33). This approach has proven to outperform Chain-of-Thought prompting for reasoning tasks by at least a 5%, while minimizing hallucinations in LLM-generated outputs (33,34). However, PS+ method comes with a downside tradeoff: it increases character usage, therefore the associated token fees. Importantly, this increased token usage may limit scalability in real-world healthcare deployment, as higher per-query costs can accumulate substantially in high-volume clinical settings and pose barriers in resource-constrained environments. Mikhail et al. showed that PS+ increased the diagnostic and next-step decision making accuracy both for DeepSeek-R1 and OpenAI o1, but it was not statically significant. Among 9 out of 10 subspecialties and overall, with and without PS+ employment, DeepSeek-R1 significantly outperformed OpenAI o1 in diagnostic and next-step decision-making.
Mikhail et al. further demonstrated that DeepSeek-R1 had lower costs per query, with cost reductions of up to 98%. This cost efficiency is particularly impactful in addressing the challenges in rural and underserved areas, where access to ophthalmology subspecialties is limited due to workforce imbalances and uneven geographic distribution (35). These shortages can translate into primary care practitioners managing ophthalmic conditions, for which they might have less experience. While per-application costs are expected to decrease with scale, developers must still account for substantial investments required for clinical trials and regulatory approvals mandated by agencies such as the FDA.
Clinical applicability should be approached with caution, as LLMs are most suitable for low-acuity triage and educational support (such as identifying cases requiring urgent specialist referral), rather than definitive diagnosis or management of high-stakes conditions where clinical expertise is essential. In this context, cost-effective tools like DeepSeek-R1 can provide real-time, guideline-based diagnostic support, enabling general practitioners to manage nonurgent cases effectively and ensure timely referrals for complex ones. Unlike proprietary models, which may impose financial burdens on healthcare systems in low- and middle-income countries through licensing and subscription fees, DeepSeek-R1 represents a viable and affordable alternative in resource-constrained settings.
While the study results are encouraging, it is important to acknowledge that ophthalmology relies heavily on pattern recognition of visual data from imaging and slit-lamp examinations (36), yet DeepSeek-R1 lacks native image analysis capabilities. Consequently, the study depended on human-generated textual descriptions of visual data, which may have introduced bias due to provider interpretation. This could have consequently affected accuracy. Similar limitations are evident in proprietary models such as OpenAI, which underperform in tasks requiring direct ophthalmologic image interpretation (30). In contrast, convolutional neural networks specifically trained for ophthalmic imaging achieve near-perfect sensitivity in detecting retinal pathology (37,38). Notably, only 19% of LLM studies in ophthalmology incorporate both textual and visual data (27). Thus, relying exclusively on text-based AI tools, as DeepSeek-R1 does, may fail to capture the level of complexity required for full comprehensive ophthalmologic care.
AI hallucinations will blur vision
One of the most significant challenges for LLMs, including advanced models like DeepSeek-R1, is their susceptibility to hallucinations characterized by fabricated responses, misinformation, and non-existent references (10,39,40). These issues undermine core ethical principles of medicine: patient safety and non-maleficence (41). While algorithmic innovations and computational advancements have helped reduce hallucination rates, these improvements remain insufficient for high-stakes healthcare applications that demand precision and accuracy and carry a risk of blindness or death if incorrectly managed. Research shows that proprietary models fabricate references nearly 50% of the time (10,42). Although DeepSeek-R1 benefits from strategies such as chain-of-thought reasoning to mitigate hallucinations, errors still occur in complex real-world scenarios (43).
Given these limitations, legal liability is a pressing concern. If LMMs provide incorrect recommendations that harm patients, it remains unclear whether legal accountability falls on developers, healthcare practitioners using the technology, institutions using the technology, or even patients themselves (17). Increased patient access to these tools amplifies the risk of misinformation, emphasizing the need for comprehensive policies and legal frameworks to mitigate such risks.
Clear vision ahead
Mikhail et al.’s work marks a significant milestone in advancing the use of open-weight LLMs in healthcare and demonstrates that these models can rival proprietary systems while offering promising cost benefits. One key to unlocking DeepSeek-R1’s full potential will be to address its lack of multimodal capability by integrating mechanisms for image interpretation. This could be achieved by merging DeepSeek-R1’s text-based reasoning abilities with powerful convolutional neural networks designed for ophthalmic image recognition. Next steps should include retrospective studies using real-world patient data from electronic health records and imaging archives, followed by prospective clinical trials in live, dynamic settings with predefined safety and efficacy endpoints. Rigorous evaluation of performance under real-world clinical conditions can help ensure that technological advances translate into meaningful improvements in patient care. Equally important, clear governance frameworks must also be established to define appropriate use, enable continuous monitoring, and address accountability and liability considerations.
In conclusion, Mikhail et al. have provided an essential and comprehensive contribution to the LLM healthcare landscape. DeepSeek-R1 has emerged as a promising open-weight alternative, delivering high-quality performance and could improve accessibility for diverse healthcare contexts for diagnostic accuracy and decision making across ophthalmology subspecialties. Despite these encouraging results, the technology remains investigational and is not yet ready for routine clinical use. Importantly, its performance does not imply equivalence to clinician judgment, and these tools should be viewed as decision-support aids rather than replacements for expert clinical decision-making.
Acknowledgments
The authors are grateful for support from the Jonas Friedenwald Professorship in Ophthalmology (to Y.M.P.) and the Wilmer Eye Institute Department of Ophthalmology.
Footnote
Provenance and Peer Review: This article was commissioned by the editorial office, Journal of Medical Artificial Intelligence. The article has undergone external peer review.
Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0009/prf
Funding: None.
Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0009/coif). Y.M.P. received support from the National Eye Institute (grant Nos. 1R01EY033000 and 1R01EY034325), the Fight for Sight-International Retinal Research Foundation (grant No. FFSGIA16002), the Alcon Research Institute Young Investigator Grant, and unrestricted departmental funding from Research to Prevent Blindness, and received consulting fee from Iridex, patents planned from University of Michigan. All these are not related to this work. The other author has no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Mikhail D, Farah A, Milad J, et al. DeepSeek-R1 vs OpenAI o1 for Ophthalmic Diagnoses and Management Plans. JAMA Ophthalmol 2025;143:834-42. [Crossref] [PubMed]
- Sallam M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare (Basel) 2023;11:887. [Crossref] [PubMed]
- Machado J. Toward a public and secure generative AI: A comparative analysis of open and closed LLMs. Int Conf Theory Pract Electron Gov ICEGOV 2025.
- Reddy K, Paulus YM. Leveraging retrieval augmented generation large language models for patient education in ophthalmology. Graefe's Archive for Clinical and Experimental Ophthalmology 2025;263:3587-90.
- Kojima T, Gu SS, Reid M, et al. Large language models are zero-shot reasoners. Adv Neural Inf Process Syst 2022;36.
- Elon University. Close encounters of the AI kind: Main report. Imagining the Digital Future. 2025. Accessed Dec 17, 2025. Available online: https://imaginingthedigitalfuture.org/reports-and-publications/close-encounters-of-the-ai-kind/close-encounters-of-the-ai-kind-main-report/
- Jay R, Davenport C, Patel R. Clinical reasoning—the essentials for teaching medical students, trainees, and non-medical healthcare professionals. Br J Hosp Med (Lond) 2024;85:1-8. [Crossref] [PubMed]
- Huesmann L, Sudacka M, Durning SJ, et al. Clinical reasoning: What do nurses, physicians, and students reason about. J Interprof Care 2023;37:990-8. [Crossref] [PubMed]
- Li DJ, Kao YC, Tsai SJ, et al. Comparing the performance of ChatGPT GPT-4, Bard, and Llama-2 in the Taiwan Psychiatric Licensing Examination and in differential diagnosis with multi-center psychiatrists. Psychiatry Clin Neurosci 2024;78:347-52. [Crossref] [PubMed]
- Busch F, Hoffmann L, Rueger C, et al. Current applications and challenges in large language models for patient care: a systematic review. Commun Med (Lond) 2025;5:26. [Crossref] [PubMed]
- Bernstein IA, Zhang YV, Govil D, et al. Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions. JAMA Netw Open 2023;6:e2330320. [Crossref] [PubMed]
- Chen JS, Reddy AJ, Al-Sharif E, et al. Analysis of ChatGPT Responses to Ophthalmic Cases: Can ChatGPT Think like an Ophthalmologist? Ophthalmol Sci 2025;5:100600. [Crossref] [PubMed]
- Gong EJ, Bang CS, Lee JJ, et al. Large Language Models in Gastroenterology: Systematic Review. J Med Internet Res 2024;26:e66648. [Crossref] [PubMed]
- Kianian R, Sun D, Crowell EL, et al. The Use of Large Language Models to Generate Education Materials about Uveitis. Ophthalmol Retina 2024;8:195-201. [Crossref] [PubMed]
- Biswas S, Logan NS, Davies LN, et al. Assessing the utility of ChatGPT as an artificial intelligence-based large language model for information to answer questions on myopia. Ophthalmic and Physiological Optics 2023;43:1562-70.
- Makary MA, Daniel M. Medical error—the third leading cause of death in the US. BMJ 2016;353:i2139. [Crossref] [PubMed]
- Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digit Med 2023;6:120. [Crossref] [PubMed]
- Marks M, Haupt CE AI. Chatbots, Health Privacy, and Challenges to HIPAA Compliance. JAMA 2023;330:309-10. [Crossref] [PubMed]
- FDA. Artificial intelligence and machine learning in software as a medical device. 2021. Accessed Dec 20, 2025. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-software-medical-device
- Wu J. The rise of DeepSeek: technology calls for the "catfish effect". J Thorac Dis 2025;17:1106-8. [Crossref] [PubMed]
- Guo D, Yang D, Zhang H, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025;645:633-8. [Crossref] [PubMed]
- Cossio M. A comprehensive taxonomy of hallucinations in large language models. arXiv 2025. arXiv:2508.01781.
- Safavi-Naini SAA, Ali S, Shahab O, et al. Benchmarking proprietary and open-source language and vision-language models for gastroenterology clinical reasoning. NPJ Digit Med 2025;8:797. [Crossref] [PubMed]
- Zhang G, Jin Q, Zhou Y, et al. Closing the gap between open source and commercial large language models for medical evidence summarization. NPJ Digital Medicine 2024;7:239.
- Lu B, Wang X, Ming W, et al. CLSeq and Nscp: novel methods for reducing hallucinations in text summarization for pre-trained models and LLMs. Journal of Big Data 2025;12:220.
- Singh S, Djalilian A, Ali MJ. ChatGPT and Ophthalmology: Exploring Its Potential with Discharge Summaries and Operative Notes. Semin Ophthalmol 2023;38:503-7. [Crossref] [PubMed]
- Zhang Z, Zhang H, Pan Z, et al. Evaluating Large Language Models in Ophthalmology: Systematic Review. J Med Internet Res 2025;27:e76947. [Crossref] [PubMed]
- Moëll B, Sand Aronsson F, Akbar S. Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1. Front Artif Intell 2025;8:1616145. [Crossref] [PubMed]
- Zou A, Wang Z, Carlini N, et al. Universal and transferable adversarial attacks on aligned language models. Int Conf Learn Represent (ICLR). 2025.
- Mihalache A, Huang RS, Cruz-Pimentel M, et al. Artificial intelligence chatbot interpretation of ophthalmic multimodal imaging cases. Eye (Lond) 2024;38:2491-3. [Crossref] [PubMed]
- Antaki F, Touma S, Milad D, et al. Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings. Ophthalmol Sci 2023;3:100324. [Crossref] [PubMed]
- Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Health 2023;2:e0000198. [Crossref] [PubMed]
- Wang L, Xu W, Lan Y, et al. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. Proc Assoc Comput Linguist Annu Meet 2023;1:2609-34.
- Wei J, Wang X, Schuurmans D, et al. Chain of thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst 2022;36.
- Resnikoff S, Lansingh VC, Washburn L, et al. Estimated number of ophthalmologists worldwide (International Council of Ophthalmology update): will we meet the needs? Br J Ophthalmol 2020;104:588-92. [Crossref] [PubMed]
- Chuck RS, Dunn SP, Flaxel CJ, et al. Comprehensive Adult Medical Eye Evaluation Preferred Practice Pattern®. Ophthalmology 2021;128:1-29. [Crossref] [PubMed]
- Anton N, Doroftei B, Curteanu S, et al. Comprehensive Review on the Use of Artificial Intelligence in Ophthalmology and Future Research Directions. Diagnostics (Basel) 2022;13:100. [Crossref] [PubMed]
- Ting DSW, Pasquale LR, Peng L, et al. Artificial intelligence and deep learning in ophthalmology. Br J Ophthalmol 2019;103:167-75. [Crossref] [PubMed]
- Hatem R, Simmons B, Thornton JE. A Call to Address AI "Hallucinations" and How Healthcare Professionals Can Mitigate Their Risks. Cureus 2023;15:e44720. [Crossref] [PubMed]
- Omar M, Sorin V, Collins JD, et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun Med (Lond) 2025;5:330. [Crossref] [PubMed]
- Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell 2023;6:1169595. [Crossref] [PubMed]
- Omar M, Nassar S, Hijazi K, et al. Generating credible referenced medical research: A comparative study of openAI's GPT-4 and Google's gemini. Comput Biol Med 2025;185:109545. [Crossref] [PubMed]
- Lin Z, Guan S, Zhang W, et al. Towards trustworthy LLMs: a review on debiasing and dehallucinating in large language models. Artificial Intelligence Review 2024;57:243.
Cite this article as: Yashar M, Paulus YM. How far can DeepSeek-R1 and OpenAI see in ophthalmology diagnosis and decision making? J Med Artif Intell 2026;9:47.

