Feasibility, safety and clinician-assessed quality of a large language model-based system for automated patient intake and clinical summarization in outpatient uveitis care
Original Article

Feasibility, safety and clinician-assessed quality of a large language model-based system for automated patient intake and clinical summarization in outpatient uveitis care

Francesco Pichi ORCID logo

Department of Ophthalmology and Vision Sciences, University of Toronto, Toronto, ON, Canada

Correspondence to: Francesco Pichi, MD. Department of Ophthalmology and Vision Sciences, University of Toronto, 340 College Street, Suite 400, Toronto, ON M5T 3A9, Canada. Email: ilmiticopicchio@gmail.com.

Background: Clinical documentation represents a major source of inefficiency in outpatient care. While artificial intelligence (AI) has advanced image-based diagnostics in ophthalmology, workflow-integrated applications targeting the documentation burden remain largely unexplored. Large language models (LLMs) offer new opportunities for automating structured data collection and clinical summarization, but the feasibility and safety of integrating such systems into subspecialty outpatient workflows have not been evaluated. The aim of this study was to evaluate the feasibility, output quality, and safety of a workflow-integrated LLM-based system for automated patient intake and clinical summarization in an outpatient uveitis clinic.

Methods: I developed and deployed a clinician-in-the-loop system in an outpatient tertiary uveitis clinic. Patients completed a structured intake questionnaire via QR code prior to the visit. Responses were processed through an orchestration layer (n8n) and summarized by a constrained LLM (GPT-4o-mini, OpenAI). Prompt constraints prohibited generation of diagnoses, treatment recommendations, or unsupported clinical inferences. Feasibility metrics included completion rate and time. Output quality was evaluated by two independent clinicians (8 and 12 years of experience) using a structured rubric assessing completeness, clinical fidelity, and usability (total score 0–6). Safety was assessed against six predefined categories of unsafe output. Inter-rater agreement was measured using Cohen’s kappa (κ).

Results: Of 62 consecutive eligible visits, 50 (80.6%) completed the questionnaire. Patients had a median age of 47 years (range, 22–74 years); 56% were female. Median completion time was 85 seconds [interquartile range (IQR), 48 seconds]. Median summary quality score was 5.0/6 (IQR, 1), with inter-rater agreement of 0.74 (κ). Overall, 88% of summaries required no or only minor edits for clinical use. No predefined unsafe outputs—including hallucinated diagnoses, medications, autonomous recommendations, or contradictions of source inputs—were identified.

Conclusions: A workflow-integrated LLM-based system for automated patient intake and clinical summarization is feasible and produces clinician-reviewable outputs under explicit safety constraints. Structured evaluation frameworks are essential for responsible deployment of such systems in clinical practice.

Keywords: Artificial intelligence (AI); clinical documentation; patient intake; large language models (LLMs); ophthalmology


Received: 20 March 2026; Accepted: 08 June 2026; Published online: 24 June 2026.

doi: 10.21037/jmai-2026-0062


Highlight box

Key findings

• A QR-based intake system using a constrained large language model (LLM) generated clinician-reviewable summaries with a median quality score of 5.0/6 and no unsafe outputs across 50 uveitis clinic visits.

What is known and what is new?

• AI in ophthalmology has focused on image-based diagnostics; documentation-oriented LLM applications in subspecialty outpatient workflows have not been systematically evaluated.

• This study provides a structured evaluation of an LLM-based documentation support system using predefined quality rubrics and explicit safety criteria in a real clinical setting.

What is the implication, and what should change now?

• LLM-based pre-visit intake can reduce redundant data collection while maintaining clinician oversight. Adoption should proceed with structured evaluation frameworks, explicit safety constraints, and attention to digital literacy barriers.


Introduction

Background

Clinical documentation remains a major source of inefficiency in outpatient care, with clinicians spending a disproportionate amount of time on data collection, synthesis, and documentation of patient-reported information (1-4). Although artificial intelligence (AI) has achieved notable successes in image-based diagnostic tasks in ophthalmology—including diabetic retinopathy screening, glaucoma detection, and retinal disease classification (5-8)—these advances have had limited impact on the documentation burden that characterizes most clinical encounters.

Large language models (LLMs)—deep neural networks trained on large text corpora to generate, summarize, and transform natural language—have shown promise in clinical documentation tasks, including discharge summary generation and clinical text summarization (9,10). However, applications targeting structured patient intake and pre-visit summarization in subspecialty outpatient settings remain largely unexplored.

Uveitis clinics provide a useful model for this problem. Clinical assessment depends not only on examination and imaging, but also on symptom evolution, systemic associations, treatment exposure, and adherence over time. Comprehensive intake is therefore essential but often consumes a substantial portion of the encounter.

Rationale and knowledge gap

Recent discussions of “agentic AI” in healthcare have emphasized that the relevant innovation is often not model autonomy, but rather orchestration of multi-step processes toward a bounded clinical goal under explicit human oversight (3,4,11). While several studies have evaluated LLMs for clinical summarization (9,10,12) and ambient documentation (13), no study has assessed the feasibility and safety of a workflow-integrated LLM system for structured patient intake and automated pre-visit summarization in an outpatient subspecialty setting.

Objective

The objective of this study was to evaluate the feasibility, output quality, and safety of a workflow-integrated LLM-based system for automated patient intake and clinical summarization in an outpatient uveitis clinic.


Methods

Study design and setting

This was a single-center implementation study conducted in an outpatient tertiary uveitis clinic between October 2025 and March 2026. The system was deployed as a clinical workflow innovation within routine care. All patients presenting for scheduled visits during the study period were eligible. Consecutive sampling was used; every eligible patient was offered the digital intake questionnaire. Available patient characteristics included age, sex, and language preference. The number of patients unable or unwilling to complete the questionnaire was recorded, along with reasons for non-completion.

System architecture

The system consisted of three integrated layers (Figure 1):

  • Patient-facing intake: patients completed a structured uveitis-specific questionnaire via QR code prior to the visit. The questionnaire captured chief complaint (CC), ocular symptoms, past ocular history (POH), systemic conditions, past medical history (PMH), review of systems (ROS), and current therapies. The full questionnaire is provided in the appendix available at https://cdn.amegroups.cn/static/public/jmai-2026-0062-1.pdf.
  • Workflow orchestration: responses were processed through an automation layer built on n8n (n8n GmbH, Berlin, Germany), an open-source workflow automation platform, which handled data ingestion, normalization, timestamping, and logging. The system was designed to operate independently of a specific electronic medical record.
  • Constrained AI summarization: a LLM [GPT-4o-mini, OpenAI, San Francisco, CA; accessed via application programming interface (API)] was used to generate a structured clinical summary using predefined headings: CC, history of present illness (HPI), ROS, POH, PMH, medications, and red flags. Prompt constraints explicitly prohibited generation of diagnoses, treatment recommendations, or unsupported clinical inferences. The model temperature was set to 0 to maximize output determinism. Outputs were limited to transformation of patient-reported inputs. The full prompt template is provided in Appendix 1.
Figure 1 Conceptual framework of an agentic AI assistant for workflow-integrated pre-visit intake in ophthalmology. (A) Traditional artificial intelligence applications in ophthalmology are predominantly image-centric and operate at a single time point, providing reactive diagnostic outputs. (B) An agentic AI assistant enables structured pre-visit data collection through QR-based patient questionnaires, workflow orchestration, and rule-constrained AI-assisted clinical synthesis, while preserving full clinician oversight. (C) The anticipated clinical impact includes reduced intake time, standardized history-taking, improved visit focus, and lower visit burden, with output delivered as a structured pre-visit clinical note (CC/HPI/POH/PMH/ROS); the approach is extendable to longitudinal monitoring and early warning applications. AI, artificial intelligence; CC, chief complaint; EMR, electronic medical records; HPI, history of present illness; LLM, large language model; OCT, optical coherence tomography; PMH, past medical history; POH, past ocular history; ROS, review of systems.

Generated summaries were reviewed by clinicians before use. No outputs were automatically entered into the medical record, and no autonomous clinical actions were triggered.

Evaluation framework

  • Feasibility: feasibility metrics included completion rate (proportion of eligible visits with fully submitted questionnaires), completion time (interval from first to last digital form timestamp), and workflow success rate (proportion of completed questionnaires processed end-to-end without system errors). Partially completed questionnaires were excluded from the analysis.
  • Summary quality: summaries were evaluated by two independent clinicians (with 8 and 12 years of ophthalmology experience, respectively) using a predefined rubric. Each summary was scored on three domains: completeness [0–2], clinical fidelity [0–2], and usability [0–2], for a total score of 0–6. Clinical fidelity was assessed by comparing each AI-generated summary against the original patient-reported questionnaire responses. Minor edits were defined as corrections of formatting, punctuation, or minor rephrasing without change in clinical content.
  • Safety: unsafe outputs were predefined as: hallucinated diagnosis, hallucinated medication, autonomous treatment recommendation, autonomous triage decision, contradiction of source input, and unsupported clinically relevant content. All summaries were systematically reviewed for these categories by both raters.

Governance and ethics

The system was clinician-in-the-loop and limited to documentation support. No diagnostic or therapeutic decisions were generated by the model. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This project was conducted as a clinical workflow innovation and quality improvement initiative and did not require Research Ethics Board approval according to institutional policy. Individual patient consent for digital intake was obtained through an on-screen consent statement prior to questionnaire completion.

Statistical analysis

Descriptive statistics were used. Continuous variables are reported as median and interquartile range (IQR), and categorical variables as counts and percentages. Inter-rater agreement was assessed using Cohen’s kappa (κ).


Results

A total of 62 consecutive visits were eligible during the study period. Of these, 50 (80.6%) completed the questionnaire and were included in the analysis. Patients had a median age of 47 years (range, 22–74 years), and 28 (56%) were female. English was the preferred language for 41 patients (82%), Mandarin for 5 (10%), and other languages for 4 (8%).

Among the 12 patients who did not complete the questionnaire, reasons included preference for verbal intake (n=5), language barriers (n=4), and technical difficulties with smartphone compatibility or connectivity (n=3).

Median questionnaire completion time was 85 seconds (IQR, 48 seconds). The workflow successfully processed 100% of completed questionnaires (50/50) end-to-end. No system-triggered clinical actions occurred.

Structured clinician review yielded a median summary quality score of 5.0/6 (IQR, 1). Inter-rater agreement was substantial (κ=0.74). Overall, 88% of summaries (44/50) required no or only minor edits (defined as corrections of formatting, punctuation, or minor rephrasing without change in clinical content) prior to clinical use. System performance metrics are summarized in Table 1 and domain-specific scores are reported in Table 2.

Table 1

System performance metrics

Metric Value
Eligible visits, n 62
Completed questionnaires, n (%) 50 (80.6)
Completion time (seconds), median [IQR] 85 [48]
Workflow success rate, n/total (%) 50/50 (100.0)
Unsafe outputs detected, n 0
Summaries requiring no/minor edits, n (%) 44 (88.0)

IQR, interquartile range.

Table 2

Summary quality scores by domain (n=50)

Domain (range, 0–2) Median [IQR] Score =2, n (%)
Completeness 2.0 [0] 42 (84.0)
Clinical fidelity 2.0 [1] 38 (76.0)
Usability 1.5 [1] 32 (64.0)
Total (range, 0–6) 5.0 [1]

IQR, interquartile range.

A de-identified example of a patient-reported questionnaire and the corresponding AI-generated summary is provided in Table S1.


Discussion

Key findings

I report the implementation and evaluation of a workflow-integrated LLM-based system designed to address a practical bottleneck in outpatient care: the collection and synthesis of patient-reported information. The system demonstrated high feasibility (80.6% completion rate, median 85-second completion time), produced clinician-reviewable summaries of substantial quality (median score 5.0/6), and generated no unsafe outputs across 50 consecutive visits. The median completion time compares favorably to the 8–10 minutes typically required for in-person history-taking in ophthalmology settings (14,15). By shifting data collection upstream and generating structured summaries before the patient enters the examination room, the system may allow clinicians to spend more of the encounter on examination and shared decision-making.

Strengths and limitations

A notable strength of this implementation is the use of a predefined, reproducible evaluation framework combining clinician scoring with explicit safety criteria. The rubric-based approach, while still limited to two raters, provides a more transparent basis for comparison than anecdotal assessment (9,10).

This study has several limitations. First, it was conducted at a single tertiary center with a limited sample of 50 visits, which may restrict generalizability. Second, the system was not integrated with the electronic medical record, requiring manual transfer of validated summaries. Third, I did not measure clinician time before and after system deployment in a controlled comparison. Fourth, patient experience, usability, and satisfaction were not formally assessed; given that the system is patient-facing, this represents a significant limitation. Fifth, patients who did not complete the questionnaire (n=12) may represent a systematically different population with lower digital literacy, introducing potential selection bias. Sixth, clinical fidelity was assessed only by clinicians comparing AI-generated summaries against questionnaire responses; patients were not involved in verifying whether their symptoms were accurately captured. Finally, the scoring rubric requires external validation across independent raters and clinical settings.

Comparison with similar research

Several recent studies have explored the use of LLMs for clinical documentation tasks. LLM-generated discharge summaries have shown clinically acceptable accuracy when compared with physician-authored notes (9,10), and ambient AI scribes have been reported to reduce documentation time while improving quality scores (13). Across multiple specialties, LLM-generated summaries have matched or exceeded human-authored notes in thoroughness, though with a tendency toward verbosity (12). My approach differs in one important respect: rather than summarizing clinician dictation or existing records, it operates upstream of the encounter, transforming patient-reported data into a structured pre-visit summary that the clinician can review before seeing the patient.

Importantly, the system was not designed to replace the clinical interview; clinicians continued to perform standard history-taking during the encounter, using the AI-generated summary as a starting point rather than a substitute. The value lies in reducing redundant data collection, not in bypassing the physician-patient interaction.

Explanations of findings

The governance model I adopted—clinician-in-the-loop, with explicit prompt constraints prohibiting diagnostic or therapeutic outputs—reflects a deliberate choice to prioritize safety over autonomy. No AI-generated summary was entered into the medical record without clinician review, and the absence of hallucinated diagnoses or medications across all 50 cases, while reassuring, should be interpreted in the context of a constrained task and a relatively small sample. As the scope of LLM applications in healthcare expands, structured safety evaluation frameworks like the one used here will be essential to distinguish systems that are genuinely safe from those that merely appear so (3,4,16-18).

The completion rate of 80.6% was reasonable for a real-world clinical setting, though lower than the >90% rates reported for electronic patient-reported outcome measures in more controlled implementations (19). The 12 non-completions—driven primarily by patient preference for verbal intake and language barriers—highlight the importance of offering alternative intake pathways alongside digital tools. The relatively younger median age of my uveitis cohort (47 years) may have contributed to the high digital engagement observed.

Implications and actions needed

The documentation burden in outpatient care has been consistently linked to clinician burnout, with physicians spending nearly twice as much time on electronic documentation and clerical tasks as on direct patient care (1,2). Pre-visit data collection systems that allow patients to provide structured information before the encounter may help reduce this burden while improving the completeness and standardization of clinical data.

Future work should evaluate longitudinal applications of this system, including integration with remote symptom monitoring for flare detection and treatment adherence tracking. Multicenter validation studies with larger and more diverse patient populations are needed to confirm generalizability. Prospective controlled studies comparing clinician documentation time with and without the system would provide stronger evidence of operational impact. Additionally, formal assessment of patient satisfaction and usability across varying levels of digital literacy will be important for equitable deployment.


Conclusions

This study demonstrates that a workflow-integrated LLM-based system for automated patient intake and clinical summarization is feasible in a real-world outpatient uveitis setting. The system produced structured pre-visit summaries of substantial quality (median 5.0/6) with no detected unsafe outputs across 50 consecutive visits. While these findings are encouraging, the limited sample size, single-center design, and absence of patient satisfaction data warrant caution in generalizing these results. LLM-based documentation support holds practical promise for reducing redundant data collection in subspecialty care, but responsible adoption requires structured evaluation frameworks, explicit safety constraints, and attention to equitable access. As clinical AI moves from proof-of-concept to deployment, the focus must shift from what these systems can do to how safely and transparently they do it.


Acknowledgments

None.


Footnote

Data Sharing Statement: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0062/dss

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0062/prf

Funding: None.

Conflicts of Interest: The author has completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0062/coif). The author has no conflicts of interest to declare.

Ethical Statement: The author is accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This project was conducted as a clinical workflow innovation and quality improvement initiative and did not require Research Ethics Board approval according to institutional policy. Individual patient consent for digital intake was obtained through an on-screen consent statement prior to questionnaire completion.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Sinsky C, Colligan L, Li L, et al. Allocation of Physician Time in Ambulatory Practice: A Time and Motion Study in 4 Specialties. Ann Intern Med 2016;165:753-60. [Crossref] [PubMed]
  2. Arndt BG, Beasley JW, Watkinson MD, et al. Tethered to the EHR: Primary Care Physician Workload Assessment Using EHR Event Log Data and Time-Motion Observations. Ann Fam Med 2017;15:419-26. [Crossref] [PubMed]
  3. Shortliffe EH, Sepúlveda MJ. Clinical Decision Support in the Era of Artificial Intelligence. JAMA 2018;320:2199-200. [Crossref] [PubMed]
  4. Bates DW, Levine D, Syrowatka A, et al. The potential of artificial intelligence to improve patient safety: a scoping review. NPJ Digit Med 2021;4:54. [Crossref] [PubMed]
  5. Ting DSW, Pasquale LR, Peng L, et al. Artificial intelligence and deep learning in ophthalmology. Br J Ophthalmol 2019;103:167-75. [Crossref] [PubMed]
  6. Schmidt-Erfurth U, Sadeghipour A, Gerendas BS, et al. Artificial intelligence in retina. Prog Retin Eye Res 2018;67:1-29. [Crossref] [PubMed]
  7. De Fauw J, Ledsam JR, Romera-Paredes B, et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat Med 2018;24:1342-50. [Crossref] [PubMed]
  8. Abràmoff MD, Lavin PT, Birch M, et al. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digit Med 2018;1:39. [Crossref] [PubMed]
  9. Patel SB, Lam K. ChatGPT: the future of discharge summaries? Lancet Digit Health 2023;5:e107-8. [Crossref] [PubMed]
  10. Rao A, Pang M, Kim J, et al. Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study. J Med Internet Res 2023;25:e48659. [Crossref] [PubMed]
  11. Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med 2019;25:44-56. [Crossref] [PubMed]
  12. Van Veen D, Van Uden C, Blankemeier L, et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat Med 2024;30:1134-42. [Crossref] [PubMed]
  13. Tierney AA, Gayre G, Hoberman B, et al. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catal Innov Care Deliv 2024; [Crossref]
  14. Baxter SL, Gali HE, Huang AE, et al. Time Requirements of Paper-Based Clinical Workflows and After-Hours Documentation in a Multispecialty Academic Ophthalmology Practice. Am J Ophthalmol 2019;206:161-7.
  15. Tai-Seale M, Olson CW, Li J, et al. Electronic Health Record Logs Indicate That Physicians Split Time Evenly Between Seeing Patients And Desktop Medicine. Health Aff (Millwood) 2017;36:655-62. [Crossref] [PubMed]
  16. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature 2023;620:172-80. [Crossref] [PubMed]
  17. Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digit Med 2023;6:120. [Crossref] [PubMed]
  18. Chen JH, Asch SM. Machine Learning and Prediction in Medicine - Beyond the Peak of Inflated Expectations. N Engl J Med 2017;376:2507-9. [Crossref] [PubMed]
  19. Schougaard LM, Larsen LP, Jessen A, et al. AmbuFlex: tele-patient-reported outcomes (telePRO) as the basis for follow-up in chronic and malignant diseases. Qual Life Res 2016;25:525-34. [Crossref] [PubMed]
doi: 10.21037/jmai-2026-0062
Cite this article as: Pichi F. Feasibility, safety and clinician-assessed quality of a large language model-based system for automated patient intake and clinical summarization in outpatient uveitis care. J Med Artif Intell 2026;9:61.

Download Citation