Co-designed, patient-facing large language models for primary-to-specialist care transitions: evidence and open questions from a pragmatic randomized trial
Most clinical evaluations of large language models (LLMs) take place in simulated settings, where the model’s output carries no downstream consequence for a real patient (1,2). Evidence from live clinical care remains scarce, and the trials that do exist are narrow: a recent scoping review found that randomized trials of artificial intelligence (AI) in clinical practice, though accumulating quickly, are dominated by single-center designs, with 63% conducted at a single site and 69% evaluating imaging-based deep-learning systems, alongside limited demographic reporting (3). The trial reported by Tao and colleagues in Nature Medicine stands out against this background because it is a pragmatic, multicenter, three-arm randomized controlled trial (RCT) of a patient-facing LLM chatbot, embedded in live outpatient workflows, with prospectively defined operational and experiential end points (4). The trial warrants a careful and balanced reading, both for what it shows and for the questions that its design leaves unanswered.
The trial and its findings
The investigators developed PreA (“pre-assessment”), a chatbot built on a commercial base model (GPT-4o mini). Rather than fine-tuning that base model on clinical data, the developers shaped its behavior through prompt augmentation and agent techniques, including a patient-facing agent built on a knowledge-graph architecture, and refined the system across a two-cycle participatory co-design process. That process involved patients, caregivers, community health workers, nurses, primary care physicians, specialists, and hospital administrators across eleven Chinese provinces. PreA pairs a patient-facing conversational interface with a clinician-facing interface, both illustrated in the original report’s interface figures. The patient-facing interface elicits history through a two-stage inquiry-and-conclusion exchange, generates one to three differential diagnoses with supporting and refuting evidence, and suggests investigations; the clinician-facing interface renders a structured referral report for the specialist to review before the face-to-face encounter (4). This distinction between a prompt-and-agent design and a fine-tuned one is central to the co-design result discussed below.
PreA was then evaluated prospectively. In the RCT, 2,138 patients at two tertiary centers in western China were randomized 1:1:1 to use PreA independently (PreA-only), to use it with staff assistance (PreA-human), or to receive usual care (No-PreA), and 2,069 were analyzed. The trial met its three primary end points. Compared with usual care, PreA-only reduced physician consultation duration from 4.41±2.77 to 3.14±2.25 minutes, a 28.7% relative reduction [95% confidence interval (CI): 22.7–34.8; P<0.001]. Physician-rated care coordination rose from 1.73±0.95 to 3.69±0.90, and patient-rated ease of communication rose from 3.44±0.97 to 3.99±0.62, both on validated 1–5 scales, corresponding to absolute increases of 1.96 and 0.55 points, which the authors express as relative increases of 113.1% and 16.0%. Secondary patient-experience measures, including attentiveness, interpersonal regard, satisfaction, and future acceptability, all favored PreA. A matched-pairs analysis comparing participating with non-participating physicians, a non-randomized contrast that the authors appropriately caveat, suggested that participating physicians saw modestly more patients per shift (28.54 vs. 24.76; a 15.3% relative increase, 95% CI: 3.4–27.2%) without longer patient waiting times. Outcomes in the PreA-only and PreA-human arms were statistically indistinguishable, which the authors interpret as evidence that the tool can operate autonomously, without staff facilitation (4).
Two further analyses extend the paper’s contribution. A prespecified examination of physicians’ clinical notes found little separation between PreA-assisted and usual-care notes, with a classifier F1 score of 0.57 and a between-condition difference below the prespecified distinguishability threshold (ΔF1 <0.02), which the authors read as reassurance against automation bias, the concern that clinicians might transcribe the model's output rather than reason independently. In a pre-trial simulation study using virtual patients, the co-designed model outperformed a counterpart that had been fine-tuned on locally collected primary care dialogues across history-taking, diagnosis, and test ordering. The data-tuned model reproduced real-world deficits, failing to offer a diagnosis in 30% of cases and to suggest indicated testing in 39% (4).
Strengths
Several features distinguish this work from the broader literature on LLMs in medicine. The trial is pragmatic and embedded, because PreA was frozen before enrollment and evaluated in routine clinical operation instead of in a sandbox, and the comparator was usual care. The sample is large, socioeconomically diverse, and drawn from the resource-limited, high-volume settings in which the global burden of fragmented referral falls hardest. The staff-assisted arm is a thoughtful design choice that probes the autonomy claim instead of assuming it. The prespecified subgroup analyses, spanning age, sex, education, income, discipline, and site, lend credibility to the consistency of the efficiency findings.
The co-design contribution is the paper’s most conceptually original claim, though the evidence for it is more preliminary than the trial’s efficacy findings and should be held to the same scrutiny. It rests not on the RCT but on the pre-trial virtual-patient simulation, the least rigorous design in the paper, in which the co-designed model outperformed a counterpart fine-tuned on locally collected primary care dialogues. With that caveat, the comparison is instructive. By setting a stakeholder-co-designed model against one fine-tuned on raw local dialogue, the investigators operationalize a question that the field has largely left implicit: should a clinical LLM mirror local practice, or help to reform it? In that simulation, passive fine-tuning on local data reproduced the same care deficits that motivated the intervention, a pattern with deep roots in the literature on algorithmic bias (5). The finding is best read as a hypothesis-generating signal that participatory design may function as a performance lever and not solely as an ethical safeguard, rather than as a settled result. It complements a companion community-codesigned primary care chatbot trial from an overlapping group (6) and suggests co-design as a transferable methodology worth testing prospectively.
Limitations
A balanced appraisal must weigh these strengths against design features that constrain what can be concluded.
The most consequential limitation is that two of the three primary outcomes rest on unblinded raters. In a pragmatic trial, patients necessarily knew whether they had used PreA, and the patient-rated experiential measures, namely ease of communication, satisfaction, and future acceptability, are the outcomes most vulnerable to expectation and Hawthorne effects. The physician-rated care-coordination outcome is similarly exposed, because specialists were masked between the PreA-only and PreA-human arms but could readily distinguish a full PreA referral from the No-PreA control, which by design contained only patient age and sex. That comparator is defensible as a representation of usual care, yet it also helps to explain the very large reported gain on this outcome, because a structured report will almost always be rated more useful than an age-and-sex stub. The size of that gain also depends on how it is expressed: reporting a rise from 1.73 to 3.69 on a five-point scale as a 113% increase inflates it through an arbitrary denominator, and the absolute shift, roughly two points for care coordination and half a point for ease of communication, is the more honest summary, a framing concern that is independent of blinding. The objectively measured 28.7% reduction in consultation duration, extracted from platform and hospital records, is the most robust of the three primary outcomes. The authors argue that this is mitigated by concordance between subjective and objective measures, by the absence of documentation differences, and by control consultation times consistent with established norms. These arguments are reasonable, but they do not exclude the possibility that some experiential gains reflect the novelty and attention of the intervention rather than the content of the referral report.
The primary efficiency end point also warrants scrutiny on its own terms. The central finding is that consultations became shorter, from roughly 4.4 to 3.1 minutes. These were already brief specialist encounters of the kind characteristic of high-volume, under-resourced health systems. A systematic review of 67 countries found that consultation times, measured most often in primary care, range internationally from under 1 minute to more than 22 minutes and are shortest in exactly such systems (7). Shorter is therefore not self-evidently better. The measured reduction is also a reduction in in-room physician consultation time and does not count the time patients spend completing the PreA pre-assessment beforehand, so the headline is partly a reallocation of effort onto patients rather than a net saving of time, a point with equity implications developed below. Most importantly, the trial measured process and experience rather than clinical substance: apart from a post hoc comparison of how closely PreA reports concorded with physician notes, it did not assess diagnostic accuracy against a reference standard, downstream utilization, repeat visits, time to correct diagnosis, or adverse events. A 28.7% reduction is a meaningful throughput gain, but whether trimming an already-compressed encounter improves diagnostic accuracy, safety, or understanding cannot be inferred from duration alone.
The quality comparison between PreA referral reports and physician notes, in which PreA scored markedly higher across history, diagnosis, and test ordering, should be read with similar caution. By the authors’ description, it pooled cases in which physician notes were absent and set a structured, purpose-built synthesis against time-pressured documentation that serves a different function. That PreA’s output reads as more complete is plausible and useful, but it does not establish that its clinical content is more correct.
The prespecified automation-bias analysis is a strength of the study, but it rests on a null result. A classifier that cannot distinguish PreA-assisted from usual-care notes is consistent with the absence of systematic influence, yet absence of detectable separation is weaker evidence than direct measurement of whether, when, and how clinicians adopted, modified, or rejected the model's diagnoses. This caution matters because contemporary LLMs do not reliably follow diagnostic or treatment guidelines and can pose real risk when their outputs are accepted uncritically (8). A tool that autonomously proposes differential diagnoses and tests to socioeconomically vulnerable patients warrant a dedicated, prospective safety end point, one that quantifies how often PreA omitted a material finding, generated an unsupported one, or anchored the physician, rather than inferring reassurance. Outright disagreement between PreA reports and physician notes was rare, at 2.8% to 5.7%, but a large share of cases could not be compared at all because physician notes were missing, so the clinical significance of any PreA-clinician divergence remains uncharacterized.
Generalizability and equity
The authors are appropriately measured about external validity, and the limits are real. The trial was conducted at two tertiary centers in western China, in Chinese-language settings, and within a self-referral system in which patients routinely bypass primary care to reach hospital specialists. Health systems with gatekeeping primary care, different documentation norms, different languages, or different liability regimes may see different effects. The model is also a moving target. The version evaluated here was frozen before enrollment, and the same architecture re-instantiated on a newer base model could behave differently, for better or worse, so any single deployment should be treated as time-locked and model-locked. Two further features of the index trial bear on how its results should be read. The tool was developed with computational resources and technical support from WeChat AI (Tencent), and PreA is described as protected intellectual property, the subject of ongoing commercial licensing discussions, and intended for development as a regulated medical device (4). None of this undermines the findings, but a commercially oriented, closed tool is harder to reproduce and to evaluate independently, and its planned path to market sharpens the equity questions that follow.
An equity tension also deserves attention, because the paper frames itself, rightly, as an equity intervention. An autonomous, patient-facing chatbot that takes the history and proposes a diagnosis is most attractive in the settings with the fewest clinicians, which are also the settings serving the most marginalized patients. This is a potential concern rather than an established consequence, but a serious one: health-information interventions with clear benefits can still widen disparities when their burdens and benefits fall unevenly, a pattern of intervention-generated inequality documented across health informatics (9). Two features of this trial illustrate the risk. First, the measured efficiency gain partly reflects effort reallocated onto patients, who must complete the PreA pre-assessment themselves, so a tool promoted on access grounds also asks under-resourced patients to do more of the work of the encounter. Second, if it were scaled without safeguards, the same efficiency logic could normalize a two-tier arrangement in which under-resourced populations interact primarily with an algorithm while better-resourced patients retain unmediated clinician time. Co-design is the best available defense against this drift, and the subgroup finding that higher-income and pediatric patients reported no significant gain in perceived physician attentiveness is a useful early-warning signal. Equity in deployment is not guaranteed by equity in design, and it should be monitored as an outcome in its own right.
Future directions
The trial points toward the studies that the field now needs. Future trials should be anchored to clinical outcomes, measuring diagnostic accuracy, appropriate test and referral selection, downstream utilization, and, above all, safety, with prospectively defined error and omission end points rather than post hoc inference. They should also adopt validated measures of patient-perceived therapeutic empathy alongside communication and satisfaction outcomes, drawing on ongoing work to standardize how empathy is defined, operationalized, and measured in AI health research, which offers a useful framework for designing such evaluations (10). Replication should extend across multiple sites, health systems, and languages, including gatekeeping systems and safety-net settings outside China, to establish where the time-reduction effect holds. Comparative work should test co-design against emerging high-quality, curated clinical dialogue datasets, because the present result argues against naive fine-tuning on raw local data, not against all data-centered strategies. Deployments should incorporate explicit guardrails for autonomous patient-facing operation, including mechanisms that surface omissions, flag low-confidence or unsupported content, and route uncertain cases to clinicians, so that autonomy is bounded by verifiable safety instead of being asserted from equivalence to a staff-assisted arm. Longer follow-up and formal cost-effectiveness analysis will be needed before throughput gains can be translated into claims about access or value.
For clinicians and health-care organizations weighing adoption now, the practical implication is narrow but usable: current evidence supports PreA-style tools for structuring and accelerating the referral hand-off and improving patients’ experience of it, but not as a substitute for clinical judgment. Early deployments are therefore best positioned as clinician-supervised aids rather than autonomous diagnostic agents, paired with local monitoring of the safety and equity signals above.
Conclusions
Tao and colleagues have produced one of the more rigorous real-world RCTs of a patient-facing clinical LLM to date. Its core contributions are a pragmatic three-arm design, a credible demonstration that the tool can operate without staff facilitation, and preliminary, simulation-based evidence that participatory co-design can outperform passive local fine-tuning. These contributions advance a field that has too often substituted benchmark performance for clinical evidence. The improvements demonstrated here, however, are in workflow efficiency and patient experience; they have not yet been shown to extend to diagnostic performance or patient safety, and several of the most striking comparisons rest on unblinded ratings or on null inference. The findings therefore justify cautious optimism, not dismissal or uncritical endorsement. PreA shows that a co-designed chatbot can streamline a referral and improve patients' experience of it, and it sets the agenda for the larger outcome-anchored and safety-anchored trials that should precede autonomous deployment at scale.
The author used a large language model (Claude, Anthropic) to assist with copy-editing, reference formatting, and consistency checking of numeric values against source documents. The author conceived and wrote the manuscript, reviewed and verified all content, and takes full responsibility for its scientific claims and interpretations; no AI tool was used to generate the analysis, arguments, or conclusions.
Footnote
Provenance and Peer Review: This article was commissioned by the editorial office, Journal of Medical Artificial Intelligence. The article has undergone external peer review.
Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0103/prf
Funding: None.
Conflicts of Interest: The author has completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0103/coif). S.B. is a paid employee of Waymark, a public benefit organization that provides free social and medical services for patients receiving Medicaid, and of HealthRight360; holds stock or stock options in Waymark and Collective Health; holds a leadership or fiduciary role at Waymark; has received consulting fees from the University of California, San Francisco; and holds two pending patents and one issued patent. No support was received for the present manuscript. The author has no other conflicts of interest to declare.
Ethical Statement: The author is accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Bedi S, Liu Y, Orr-Ewing L, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA 2025;333:319-28. [Crossref] [PubMed]
- Goh E, Gallo RJ, Strong E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat Med 2025;31:1233-8. [Crossref] [PubMed]
- Han R, Acosta JN, Shakeri Z, et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit Health 2024;6:e367-73. [Crossref] [PubMed]
- Tao X, Zhou S, Ding K, et al. An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial. Nat Med 2026;32:934-42. [Crossref] [PubMed]
- Obermeyer Z, Powers B, Vogeli C, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019;366:447-53. [Crossref] [PubMed]
- Li S, Li Y, Zhou S, et al. A community-codesigned LLM-powered chatbot for primary care: a randomized controlled trial. Nat Health 2026;1:238-50. [Crossref] [PubMed]
- Irving G, Neves AL, Dambha-Miller H, et al. International variations in primary care physician consultation time: a systematic review of 67 countries. BMJ Open 2017;7:e017902. [Crossref] [PubMed]
- Hager P, Jungmann F, Holland R, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med 2024;30:2613-22. [Crossref] [PubMed]
- Veinot TC, Mitchell H, Ancker JS. Good intentions are not enough: how informatics interventions can worsen inequality. J Am Med Inform Assoc 2018;25:1080-8. [Crossref] [PubMed]
- Howcroft A, Benford S, Valstar M, et al. Empathy in AI for Health and Care Settings-Definition, Expression, and Measurement: Protocol for a Scoping Review. JMIR Res Protoc 2026;15:e93078. [Crossref] [PubMed]
Cite this article as: Basu S. Co-designed, patient-facing large language models for primary-to-specialist care transitions: evidence and open questions from a pragmatic randomized trial. J Med Artif Intell 2026;09:74.

