A real-time artificial intelligence-integrated system (DrAidTM Endo) feasibility in identifying anatomical landmarks and detecting upper gastrointestinal tract lesions: a randomized controlled trial
Highlight box
Key findings
• This clinical trial evaluated the feasibility of DrAidTM Endo, an artificial intelligence (AI)-integrated system that assists endoscopists in identifying anatomical landmarks and detecting 5 type lesions of upper gastrointestinal (GI). The accuracy of DrAidTM in detecting reflux esophagus, duodenal ulcer and gastritis was 81.4%, 49.0% and 33.3% respectively. Accuracy in detecting malignant lesions needs to be proven by further research.
• The accuracy of AI-assisted system in identifying 10 anatomical landmarks ranged from 91.2% to 100%.
What is known and what is new?
• AI is playing an increasingly important role in endoscopy, with notable progress in colonoscopy applications. However, its effectiveness in upper gastrointestinal endoscopy, especially in developing countries, remains limited. The system that integrates multiple algorithms to detect diverse types of lesions is not yet widely adopted.
• To our knowledge, this study was the first one to assess the real-time performance of AI models detecting multiple lesions in upper gastrointestinal endoscopy.
What is the implication, and what should change now?
• Based on our findings, further refinements are needed to enhance real-time performance of the AI products in detecting duodenal ulcer and gastritis.
• Multicenter clinical trials with larger sample sizes are needed, focusing on the AI’s ability to support the detection of specific lesions, especially malignant lesions, to further assess the AI’s performance.
Introduction
Upper gastrointestinal (GI) endoscopy plays an important role in timely detection and treatment of GI conditions, especially malignancy (1-4). However, the quality of endoscopic procedures depends heavily on equipment, preparation, and endoscopist’s expertise (5-8). Despite available guidelines and consensus for upper GI endoscopy, poor adherence and lack of proper supervision are common in developing countries (9).
A remarkable proportion of GI lesions is missed in endoscopy [42% for chronic gastritis (10), 6% to 14% for upper GI cancer (8,11-14) and up to 50% for Barrett’s esophagus (15)]. In a study on patients with undetected upper GI cancer lesions, 73% of these missed lesions were visible during the first endoscopy and could have been treated (16). In another study on Vietnamese patients with advanced gastric cancer (GC), around 65% had previously undergone endoscopy, mostly within 2 years before receiving the diagnosis of cancer (17). These findings emphasize the need for additional strategies to improve the quality of endoscopic procedures and reduce the missing rates.
Artificial intelligence (AI) has been researched as a supportive tool for GI endoscopy (18). Many AI algorithms, either deployed on a device/platform or still in the refinement phase, have focused on detecting or classifying a single lesion such as esophageal cancer (EC), Barrett’s esophagus or GC. However, to meet the demand of endoscopists, an “all-in-one” system should be capable of efficiently performing on multiple lesions. Since 2019, research groups in Vietnam have developed algorithms for multiple lesions and validated them using still images (19-22). However, the effectiveness of AI-assisted upper GI endoscopy has not been confirmed on videos and in real time.
In 2024, the Institute of Gastroenterology and Hepatology (IGH) and NVIDIA Vietnam developed DrAidTM Endo, an AI-integrated system that assists endoscopists in real-time detection of five types of upper GI lesions including reflux esophagitis (RE), gastritis, duodenal ulcer (DU), EC and GC. In this present article, we report the results of a randomized clinical trial conducted to (I) evaluate the effectiveness of DrAidTM Endo by comparing the real-time detection rate of five types of upper GI lesions with conventional endoscopy and (II) determine the predictive accuracy of DrAidTM Endo in detecting and classifying upper GI lesions during endoscopy. We also investigated the accuracy in identifying anatomical landmarks. We present this article in accordance with the CONSORT-AI reporting checklist (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-206/rc).
Methods
Study design
We conducted a randomized controlled clinical trial, single-blind at IGH and Hoang Long Clinic from January 2024 to December 2024. Patient recruitment took place from June 2024 to September 2024.
Study population
Eligible participants must be 18 years or older, have an indication for upper GI endoscopy (cancer screening, surveillance of pre-malignant conditions such as atrophic gastritis or post-ulcer treatment, or the presence of clinical symptoms including epigastric pain, dyspepsia, dysphagia, or anemia of unknown origin), and have provided informed consent. Exclusion criteria include: (I) history of esophageal, gastric, or intestinal resection surgery; (II) requiring observation or procedures that consume more than 25% of the total examination time; (III) safety concerns related to anesthesia or endoscopy procedures such as high bleeding risk or severe comorbid diseases; (IV) emergency intervention during endoscopy; (V) excessive gastric fluid or food residue obstructs visualization [Polprep: Effective Assessment of Cleanliness in Esophagogastroduodenoscopy (PEACE) score <2] (23).
Intervention
DrAidTM Endo is an AI-integrated system developed by IGH in collaboration with NVIDIA Vietnam, which can be used on endoscopy systems that can output video signals via high-definition multimedia interface (HDMI) or digital visual interface (DVI)-out ports. Its main functions include identifying anatomical landmarks and detecting and classifying GI lesions in real-time. The AI model integrated into the system was developed based on two models: EfficientNet B5 for anatomical landmark identification and YOLOv8 for lesion detection and classification.
Anatomical landmark identification
The EfficientNet B5 model (24,25), consisting of 30 million parameters, was pre-trained on the ImageNet database (26). The model was further trained on 44,337 high-resolution endoscopic images of the upper GI tract and validated on 2,287 still images. Training images were 224×224 pixels, red, green and blue (RGB)-colored, and properly labelled with one of ten anatomical landmarks by experienced endoscopists. The landmark labels included 10 main anatomical landmarks: (I) pharynx, (II) esophagus, (III) cardia, (IV) gastric body, (V) fundus, (VI) antrum, (VII) greater curvature, (VIII) lesser curvature, (IX) duodenal bulb, (X) duodenum, and a separate label “other region” (which are images difficult to classify due to insufficient inflation/too close to the mucosa, images with poor quality such as mucus, glare, too dark, blurry, shaky, bubbles, etc.). To improve dataset diversity, image augmentation techniques were also applied. All training images were preprocessed and standardized prior to training. The model was trained for 30 epochs and achieved an average accuracy of 95.2% in detecting ten anatomical landmarks (Figure S1).
Lesion detection and classification
The YOLOv8 model (27) was pre-trained on the Common Objects in Context (COCO) dataset (28). It was further trained on 12,002 endoscopic images, tested on 1,340 images and validated in 47 videos. Training images were 640×640 pixels, RGB-colored and properly delineated for lesions within the study scope using boundary coordinates (manually labeled by 12 endoscopists ≥3 years of experiences then verified by 7 experienced endoscopists with ≥5 years of experience; the disagreement between experts was then discussed to reach final consensus). Delineated lesions were classified into one of five categories: RE, gastritis, DU, EC, and GC. The model was trained for 500 epochs. The training sensitivity and specificity for video data were 91.1% vs. 92.6% for RE, 95.6% vs. 95.9% for EC, 65.3% vs. 73.9% for gastritis, 91.4% vs. 94.9% for GC, and 92.3% vs. 94.6% for DU, respectively. The AI-generated bounding boxes underwent post-processing to eliminate false-positive predictions (Figure S2).
The AI models were then integrated into a compact computer running the Ubuntu operating system and this device can be connected to the output ports (HDMI or DVI-out) of the endoscopy system to collect video data. The technical specifications of this system was presented in the Table S1. In the real-time mode, DrAidTM Endo simultaneously extracts and analyzes the video data at 42fps, provides predictions in landmark identification (excluding label “others”) and lesion detection/classification, and displays on a separate monitor (minimum resolution of 720×1,280 pixels) the real-time video with prediction results with a low latency (150 milliseconds) (Figure 1). Endoscopists can use AI’s real-time predictions as suspected lesions to examine more thoroughly and make the final diagnosis.
All the image datasets used for training and validation was completely independent from the patient data used in this clinical trial.
Sample size
We calculated the sample size to adequately power the comparison of the real-time detection rate of one of our five GI lesions. In literature review, detection rates without AI-assisted system ranged from 2% to 25% (29-33); therefore, we chose an average detection rate of 6.2%. We expected that using AI-assisted system would increase the detection rate by 10% (from 6.2% to 16.2%). With a two-sided α level of 0.05 and power of 80%, the minimum sample size needed to detect such a statistically significant difference was 154 (77 patients in each group).
Study procedures
All patients indicated for upper GI endoscopy underwent a screening examination by a physician. Eligible patients provided informed consent were randomly allocated to the intervention group (conventional endoscopy with AI-assisted system) or the control group (conventional endoscopy) by using sealed envelope. The random sequence was generated using the website “https://www.sealedenvelope.com” with a 1:1 ratio (block size of 4). Participants were blinded during study process.
Endoscopic procedures were performed using the FujiFilm 7000 endoscopic system under white light imaging (WLI) and three image-enhanced endoscopy (IEE) modes including blue light imaging (BLI), linked color imaging (LCI) and flexible spectral imaging color enhancement (FICE). Participating endoscopists must have at least 5 years of experience, have performed at least 1,000 endoscopic procedures, and are proficient at interpreting and evaluating endoscopic images of various anatomical locations and lighting modes. Before study implementation, all endoscopists received standardized training on the AI-assisted system, 10 anatomical landmarks and five lesion classification criteria used in the study.
During endoscopy, collected information included mucosal cleanliness, anatomical landmarks, detected lesions (location, number, size, and classification). The cleanliness of gastric mucosal was classified using a four-level PEACE scale (23,34). Identification of anatomical landmarks (ten locations used for the AI model) was based on the European Society of Gastrointestinal Endoscopy (ESGE) guideline (35). RE was categorized into five subtypes according to the modified Los Angeles classification (36,37). Gastritis was divided into six subtypes based on the Sydney and Kyoto classifications (38,39) including the following: nodularity, mucosal atrophy, raised erosion, flat erosion, hematin, and spotty redness. DU included six subtypes based on the Sakita and Fukutomi classification (40). EC and GC each had six subtypes according to the respective Japan Esophageal Society and Japanese Gastric Cancer Association’s classifications (41,42). For suspected malignancies, histopathological confirmation will be performed by expert pathologists following the 2019 World Health Organization (WHO) classification for esophageal and GC (43). Examination time was measured using a stopwatch. We excluded time spent on additional procedures such as polypectomy, Helicobacter pylori (H. pylori) testing, or biopsy sampling. In both study groups, the required minimum examination time was seven minutes, with at least four minutes dedicated to gastric examination (44,45).
The entire endoscopic process, including the AI-assisted detection and classification, was recorded for further evaluation. An expert endoscopist with at least 10 years of experience (who already received standardized training before the study) later reviewed and analyzed the recorded videos with AI predictions to evaluate the real-time performance of AI-assisted system for correct anatomical identification and lesion detection/classification. This assessment was directly based on the primary diagnosis of performing endoscopists. In case of any discrepancy reported during this process between the primary diagnoses of endoscopists and secondary reviews of the expert, the two physicians will discuss determining the result.
Statistical analysis
Data were collected in a case report form, entered into REDCap, and cleaned and analyzed using R using the per-protocol approach because the number of excluded cases was low, the exclusion was likely due to chance, and it was difficult to handle the missing outcomes using imputation. Continuous variables were presented as mean (± standard deviation) and compared using the t-test. Ordinal and categorical variables were presented as counts (percentage) and compared using the Chi-squared test or Fisher’s exact test, where appropriate. A P value <0.05 were considered statistically significant.
The primary outcome of this trial was the real-time detection rate of five types of upper GI lesions (RE, gastritis, DU, EC, and GC) calculated by the proportion of patients diagnosed with each lesion type among the total number of patients recruited in each group.
The system’s performance was evaluated by sensitivity, specificity, accuracy, and F1-score for correct lesion detection and classification by patient with the gold standard being the final expert diagnosis. Additionally, the gold standard for malignant cases (EC, GC) was histopathology confirmation of cancer. A true positive (TP) case was defined as one in which the patient had lesion X on endoscopy, and the AI-assisted system detected at least one region of lesion X and correctly classified that region as lesion X during endoscopy. A false negative (FN) case was defined as one in which the patient had lesion X but the AI-assisted system either failed to detect any region of lesion X or failed to classify any correctly detected region as lesion X. A false positive (FP) case occurred when patients did not have lesion X but the AI-assisted system incorrectly detected and classified at least one region as lesion X. A true negative (TN) case occurred when the patients did not have lesion X and the AI-assisted system did not classify any detected region as lesion X (see Figure S3 for more details). Sensitivity is defined as TP divided by the total number of positives: sensitivity = TP/(TP + FN). Specificity is defined as TN divided by the total number of negatives: specificity = TN/(TN + FP). Accuracy is defined as (TP + TN)/(TP + TN + FP + FN). F1-score is defined as (2 × TP)/(2 × TP + FP + FN).
Ethics
The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study has been approved by the Institutional Ethics Review Board of Hanoi Medical University (IRB-VN01.001/IRB00003121/FWA 00004148) under approval No. 1502/GCN-HDDCYSH-DHYHN. Informed consent was obtained for all participants.
Results
Patients’ characteristics
Between June 2024 and September 2024, we screened 237 patients, of whom 218 agreed to participate in the study. Twelve patients were excluded after randomization due to safety concerns related to anesthesia, and one patient withdrew from the study. Also, four patients in the intervention group were switched to the control group due to AI-assisted system malfunction, one patient was switched from control group to intervention group due to note-taking errors by research coordinators and endoscopy technicians. A total of 205 patients (102 in the intervention group and 103 in the control group) were included in the data analysis. The study flow is summarized in the CONSORT flow diagram in Figure 2.
The mean age of the participants was 45.1±11.5 years, with 61.5% female. More than three-quarter (77.6%) of patients presented with at least one upper GI symptom. In which, epigastric pain (37.6%), bloating (19.5%), and belching (17.6%) were among the most frequently reported symptoms (Table 1).
Table 1
| Characteristics | Intervention (n=102) | Control (n=103) | Total (n=205) | P value |
|---|---|---|---|---|
| Female sex | 62 (60.8) | 64 (62.1) | 126 (61.5) | 0.96a |
| Age (years) | 45.3 [11.9] | 44.8 [11.2] | 45.1 [11.5] | 0.78b |
| Upper GI symptoms | 81 (79.4) | 78 (75.7) | 159 (77.6) | 0.64a |
| Epigastric pain | 43 (42.2) | 34 (33.0) | 77 (37.6) | 0.23a |
| Bloating | 19 (18.6) | 21 (20.4) | 40 (19.5) | 0.89a |
| Belching | 19 (18.6) | 17 (16.5) | 36 (17.6) | 0.83a |
| Regurgitation | 18 (17.6) | 15 (14.6) | 33 (16.1) | 0.17a |
| Acid reflux | 7 (6.9) | 14 (13.6) | 21 (10.2) | 0.68a |
| Other GI symptoms | 24 (23.5) | 29 (28.2) | 53 (25.9) | 0.55a |
Data are presented as n (%) or mean [SD]. a, Pearson’s Chi-squared, b, t-test. GI, gastrointestinal; SD, standard deviation.
Detection and classification of GI lesions
Table 2 presents the upper GI endoscopy results of both groups. There were no significant differences between the two groups regarding endoscopic lesion detections, examination time, or mucosal cleanliness (P>0.05). Compared to the control group, the intervention group had higher detection rates in RE (78.4% vs. 76.7%), lower in gastritis (90.2% vs. 92.2%) and DU (8.8% vs. 16.5%). Additionally, the average examination time was longer in the intervention group (422.8±40.8 seconds) than in the control group (417.6±36.8 seconds). All differences were not statistically significant (Table 2).
Table 2
| Lesions | Intervention (n=102) | Control (n=103) | Total (n=205) | P value |
|---|---|---|---|---|
| Reflux esophagitis | 80 (78.4) | 79 (76.7) | 159 (77.6) | 0.90a |
| Esophageal cancer | 2 (2.0) | 0 | 2 (1.0) | – |
| Gastritis | 100 (98.0) | 102 (99.0) | 202 (98.5) | 0.99a |
| Gastritis† | 92 (90.2) | 95 (92.2) | 187 (91.2) | 0.79a |
| Gastric cancer | 2 (2.0) | 0 | 2 (1.0) | – |
| Duodenal ulcer | 9 (8.8) | 17 (16.5) | 26 (12.7) | 0.15a |
| Total mucosal cleanliness score‡ | 8.96 [0.2] | 9.00 [0.0] | 8.98 [0.2] | 0.10b |
| Examination times | 422.8 [40.8] | 417.6 [36.8] | 420.2 [38.8] | 0.35b |
Data are presented as n (%) or mean [SD]. †, only includes six subtypes of gastritis previously used to train the AI models: nodularity, mucosal atrophy, raised erosion, flat erosion, hematin, and spotty redness; ‡, cleanliness is assessed once pumping and cleaning are complete; a, Pearson’s Chi-squared, b, t-test. GI, gastrointestinal; SD, standard deviation.
Out of the 831 lesions identified (excluding cancer lesions), gastritis and RE were the most common lesions, comprising 61.3% and 33.2%, respectively. No significant differences were observed between the control and intervention groups in terms of lesion count, size, or classification of RE, gastritis, and DU lesions (P>0.05) (Table S2). Specifically, Los Angeles Classification (LA) grade A was the most common RE lesion (88.7%). Common gastritis findings included flat erosions (26.5%), spotty redness (29.9%), and mucosal atrophy (14.3%). Most localized lesions measured <5 mm (46.2%). Regarding DU, the main lesion was superficial DU, <5 mm lesions more frequent in the intervention group (96.9%) than the control group (56%) (Table S2). Four patients had confirmed malignant lesions (GC: type 2 and 4; 2 EC: type 0), all exceeding 10 mm (Figure 3).
DrAidTM Endo demonstrated the highest accuracy in detecting RE (81.4%) while lowest in detecting gastritis (33.3%) (Table 3 and Table S3). Although DrAidTM Endo successfully detected all DU lesions, its specificity (44.1%) and accuracy (49.0%) in detecting DU were lower than in detecting RE. Regarding EC and GC, the AI-assisted system correctly detected 3/4 cancer cases and misidentified an early-stage EC lesion (20 mm × 30 mm) (Figure 3).
Table 3
| Index | RE | Gastritis | DU |
|---|---|---|---|
| Sensitivity (95% CI), % | 77.5 (66.8–86.1) | 28.3 (19.4–38.6) | 100.0 (66.4–100.0) |
| Specificity (95% CI), % | 95.5 (77.2–99.9) | 80.0 (44.4–97.5) | 44.1 (33.7–54.7) |
| Accuracy (95% CI), % | 81.4 (72.4–88.4) | 33.3 (24.3–43.4) | 49.0 (39.0–59.1) |
| F1-score (95% CI)†, % | 86.7 (80.0–92.1) | 43.3 (31.6–53.9) | 25.7 (12.3–39.4) |
†, the 95% CI for the F1-score of each lesion was estimated using bootstrap resampling with 10,000 iterations. AI, artificial intelligence; CI, confidence interval; DU, duodenal ulcer; RE, reflux esophagitis.
During real-time processing, the DrAidTM Endo simultaneously reviewed and delineated regions on the video frame for lesion detection. Therefore, in some cases, frames can contain misdetections. These errors occurred in 4.9% of cases for RE, 35.5% of cases for gastritis, and 52.9% of cases for DU (n=102). In total, 121 incorrectly detected regions were reported by the AI-assisted system. Most misdetections occurred in areas affected by blur/reflections (40.5%) and blood-stained areas (20.7%). Among these, 58.1% of blur/reflection-related errors were classified as DU, while 54.8% of blood-stained regions, primarily at rapid urease test sites, were misclassified as gastritis (Table S4). Misdetections were most frequent in the duodenal bulb (42.1%), the antrum (28.9%), and the duodenum (10.7%). Regarding image lighting modes, misdetection was more common under IEE modes, particularly BLI (31.4%) and LCI (30.6%).
Anatomical landmark identification accuracy
DrAidTM Endo correctly identified anatomical landmarks with a mean accuracy of 98.8%; the lowest accuracy was in pharynx (91.2%). During processing, DrAidTM Endo still misidentified certain frames despite correctly recognizing the anatomical location in most other frames. The most frequently reported misidentified locations were the antrum, esophagus, pharynx, and fundus (Figure 4).
Discussion
To enhance procedure quality and meet the growing demand for lesion detection in Vietnam, we conducted a clinical trial study to evaluate the effectiveness of an AI-integrated system (DrAidTM Endo), a real-time AI-assisted system designed to assist doctors in identifying 10 anatomical landmarks and detecting/classifying five types of lesions during upper GI endoscopy. We found that DrAidTM Endo demonstrated an overall accuracy of 98.8% in anatomical landmark identification, achieved the highest accuracy of 81.4% in RE among the three common lesions at and is capable of detecting malignant lesions. There were no significant differences between the two groups regarding endoscopic lesion detections.
According to ESGE guidelines, 10 anatomical landmarks must be identified in each upper GI endoscopy session to improve lesion detection (44). According to Huang et al. [2015], over 68% of early GC cases were overlooked in the lesser curvature because this area was not adequately examined during endoscopy (46). This finding underscores the importance of thoroughly observing the entire anatomical region, while AI offers the potential to reduce such omissions during the procedure. DrAidTM Endo correctly identified 98.82% of all anatomical landmarks, with over 95% accuracy for most landmarks using the EfficientNet B5 model. One of the biggest advantages of EfficientNet B5 over previous models is its low parameter count (30 million), leading to fast computation speed and minimal memory usage while maintaining a high accuracy (24). Previous studies have also explored AI’s ability to recognize anatomical landmarks to improve endoscopic procedures’ quality. A convolutional neural network (CNN) algorithm developed by Choi et al. [2022], trained on 2,599 endoscopic images, achieved 97.58% accuracy in classifying 8 anatomical landmarks (47). The WISENSE model, built on the Visual Geometry Group-16 (VGG-16) algorithm, was trained to identify 26 anatomical landmarks and blind spots in real-time. Clinical trials demonstrated that the system significantly reduced blind spots in upper GI endoscopy compared to the control group (5.9% vs. 22.5%), with over 90% accuracy in identifying blind spots in the antrum. The model achieved an overall accuracy of 90.2% (48).
In our study, the lesion detection rates were comparable between the AI-assisted group and the control group. This may be explained by several factors. First, our participating endoscopists were experienced physicians, who would be able to accurately detect and classify diverse and complex lesions in the upper GI tract with low missing rates in both groups. Second, the required minimum endoscopy time was 7 minutes, giving endoscopists adequate time to conduct comprehensive mucosal examination, leaving no blind spots and thus minimizing the risk of missed lesions. Last, given the heterogeneity of lesion classifications, our sample size did not provide adequate power to evaluate the impact of AI-assisted system for each lesion type.
DrAidTM Endo achieved 77.5% sensitivity and 95.5% specificity in diagnosing RE with the lowest misdetection rate (4.9%). Real-time AI applications for detecting RE remain limited. Ge et al. used an EfficientNet-based model to categorize RE lesions into 3 groups (normal, LA A+B, and LA C+D), achieving 95.7% of accuracy on external datasets, though it lacks real-time validation (49). A clinical trial using a ResNet-50-based model for diagnosis of gastroesophageal reflux disease (GERD) in patients with reflux symptoms reported a sensitivity of 88.9% and a specificity of 58.5% in 78 patients (50). While our system showed lower sensitivity, this may be due to differences in training data. The ResNet-50 model used IEE images under magnification focused on intrapapillary capillary loops (IPCLs), a typical finding of GERD, whereas our model was trained on non-magnified IEE images.
Previous studies on AI-assisted gastritis detection in endoscopy primarily focused on atrophic gastritis, intestinal dysplasia, and H. pylori detection, mostly on still images (51-54). DrAidTM Endo provide real-time detection that targets six gastritis-related lesions (nodularity, mucosal atrophy, raised erosion, flat erosion, hematin, and spotty redness) based on the Sydney and Kyoto classifications. However, its sensitivity in our study was low (28.3%). Similarly, low accuracy was reported in a clinical trial using the ENDOANGEL model, with a detection rate of 7.1% for non-biopsy gastric lesions (n=437). Specifically, the highest positive prediction rate was 16.7% for hemorrhagic erosion (n=6), 15.8% for raised erosion (n=19), and 11.7% for flat erosion (n=163) (55). One possible reason for the low detection rate in our study is that many lesions were smaller than 5 mm (43.1%) or had polymorphic morphology (41.6%), making it difficult to delineate lesion boundaries and hinder effective AI training.
Regarding DU, DrAidTM Endo achieved a sensitivity of 100% but accuracy of 49% due to a high number of FPs. Most false detections of DU resulted from blurring or reflections, due to the similarities between light reflections on the mucosal surface and the characteristics of DU lesions, which typically appear white and brighter than other lesion types. Additionally, this issue was higher under BLI mode, where contrast enhancement makes reflective areas resemble small ulcerative lesions. This study also provides valuable data in real-time setting and highlights the need for further approaches, training strategies, and optimizing AI algorithms on gastritis and DU lesions.
Misdetections are technical drawbacks of AI detection algorithms due to processing a large number of video frames during real-time application (56). In this study, most of non-lesion regions detected by AI were blur/reflective areas, normal mucosal, and blood areas (post-H. pylori test hemorrhagic regions). Although experienced endoscopists can easily differentiate misdetections, these are still inconvenient and can lead to alert fatigue (57). A major cause of misdetections is poor-quality frames in real-world endoscopy, such as overexposure, blurriness, or bubbles, whereas the training dataset primarily consists of high-quality images. Therefore, further training with a broader range of endoscopic image’s quality, especially video-based images, may be a good approach to reduce these errors (58).
To date, no studies have assessed the real-time performance of more than five AI models running simultaneously, as demonstrated in our study (one model for anatomical landmarks, five models for lesion detection/classification). While our findings suggested that this multi-model approach may affect overall system accuracy, it demonstrates technical feasibility and potential for further development in clinical settings. Studies indicate that endoscopists generally agree AI should act as a secondary observer, functioning as a co-supervisor and supporting the training of novice endoscopists (59,60). ESGE guidelines also emphasize AI’s role in enhancing performance among less experienced endoscopists and ensuring consistent quality in the diagnosis and treatment of GI neoplasm (61). The integration of AI in endoscopy training may help bridge the skill gap between novice and expert endoscopists (61,62).
This study has some limitations, including a small patient cohort, data collected from a single endoscopy center, and involvement of only three endoscopists. Since all cancer cases were in the intervention group, we cannot compare the performance in detecting cancer between the two group. Additionally, our AI algorithms for detection of gastritis were trained primarily on localized lesions (flat erosions, raised erosions, hematin), which may limit their performance since gastritis morphology is diverse and complex. In the future, to achieve more distinct results, we aim to optimize the AI-assisted system by expanding the dataset, incorporating lower GI endoscopy, and conducting further studies to include less-experienced endoscopists and a larger sample size in multi-center, focusing on the AI’s ability to support the detection of specific lesions, especially malignant lesions, to develop an “all-in-one” real-time AI-assisted endoscopy tool.
Conclusions
In conclusion, no significant difference was observed in the real-time detection rate of five types of upper GI lesions between the AI-assisted group and the control group. The DrAidTM Endo system showed high accuracy in identifying anatomical landmarks and in detecting and classifying RE. Further optimization is required to improve the system’s performance in detecting and classifying DU and gastritis. Future research should focus on developing and optimizing an “all-in-one” real-time AI-assisted endoscopy tool by expanding the dataset, incorporating lower GI endoscopy, and conducting multi-center studies.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the CONSORT-AI reporting checklist. Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-206/rc
Trial Protocol: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-206/tp
Data Sharing Statement: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-206/dss
Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-206/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2025-206/coif). H.D.V. received honoraria for lectures, presentations, educational events from Abbott, Astrazeneca, Zuellig, Johnson and Johnson, Biocodex, and participation on an Advisory Board for Gilead All4Liver Grant and Gilead Global Public Health Grant. She is Vice General Secretary of the Vietnam Association of Gastroenterology, Vice General Secretary of the Vietnam Association for the Study of Liver Diseases, Vice President of the Vietnam Youth Federation and President of the Global Vietnamese Youth and Student Association. H.N.N., G.D.Q., D.T.A., H.N.V., K.D.T., and M.T.N. are employee of VinBrain Joint Stock Company (Hanoi, Vietnam) and NVIDIA Corporation (Santa Clara, US). S.T.Q.H. is CEO of VinBrain Joint Stock Company (Hanoi, Vietnam) and Vice President of NVIDIA Corporation (Santa Clara, USA). L.D.V. is Permanent Vice President of Vietnam Association of Gastroenterology (VNAGE). The other authors have no other conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study has been approved by the Institutional Ethics Review Board of Hanoi Medical University (IRB-VN01.001/IRB00003121/FWA 00004148) under approval No. 1502/GCN-HDDCYSH-DHYHN. Informed consent was obtained for all participants.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Spiceland CM, Lodhia N. Endoscopy in inflammatory bowel disease: Role in diagnosis, management, and treatment. World J Gastroenterol 2018;24:4014-20. [Crossref] [PubMed]
- Săftoiu A, Hassan C, Areia M, et al. Role of gastrointestinal endoscopy in the screening of digestive tract cancers in Europe: European Society of Gastrointestinal Endoscopy (ESGE) Position Statement. Endoscopy 2020;52:293-304. [Crossref] [PubMed]
- Kuribayashi S, Hosaka H, Nakamura F, et al. The role of endoscopy in the management of gastroesophageal reflux disease. DEN Open 2022;2:e86. [Crossref] [PubMed]
- Katai H, Ishikawa T, Akazawa K, et al. Five-year survival analysis of surgically resected gastric cancer cases in Japan: a retrospective analysis of more than 100,000 patients from the nationwide registry of the Japanese Gastric Cancer Association (2001-2007). Gastric Cancer 2018;21:144-54. [Crossref] [PubMed]
- Zhao S, Wang S, Pan P, et al. Magnitude, Risk Factors, and Factors Associated With Adenoma Miss Rate of Tandem Colonoscopy: A Systematic Review and Meta-analysis. Gastroenterology 2019;156:1661-1674.e11. [Crossref] [PubMed]
- Yoshimizu S, Hirasawa T, Horiuchi Y, et al. Differences in upper gastrointestinal neoplasm detection rates based on inspection time and esophagogastroduodenoscopy training. Endosc Int Open 2018;6:E1190-7. [Crossref] [PubMed]
- Lee GJ, Park SJ, Kim SJ, et al. Effectiveness of Premedication with Pronase for Visualization of the Mucosa during Endoscopy: A Randomized, Controlled Trial. Clin Endosc 2012;45:161-4. [Crossref] [PubMed]
- Menon S, Trudgill N. How commonly is upper gastrointestinal cancer missed at endoscopy? A meta-analysis. Endosc Int Open 2014;2:E46-50. [Crossref] [PubMed]
- Rutter MD, Rees CJ. Quality in gastrointestinal endoscopy. Endoscopy 2014;46:526-8. [Crossref] [PubMed]
- Du Y, Bai Y, Xie P, et al. Chronic gastritis in China: a national multi-center survey. BMC Gastroenterol 2014;14:21. [Crossref] [PubMed]
- Chadwick G, Groene O, Hoare J, et al. A population-based, retrospective, cohort study of esophageal cancer missed at endoscopy. Endoscopy 2014;46:553-60. [Crossref] [PubMed]
- Amin A, Gilmour H, Graham L, et al. Gastric adenocarcinoma missed at endoscopy. J R Coll Surg Edinb 2002;47:681-4.
- Delgado Guillena PG, Morales Alvarado VJ, Jimeno Ramiro M, et al. Gastric cancer missed at esophagogastroduodenoscopy in a well-defined Spanish population. Dig Liver Dis 2019;51:1123-9. [Crossref] [PubMed]
- Kim HH, Cho EJ, Noh E, et al. Missed synchronous gastric neoplasm with endoscopic submucosal dissection for gastric neoplasm: experience in our hospital. Dig Endosc 2013;25:32-8. [Crossref] [PubMed]
- Singer ME, Odze RD. High rate of missed Barrett's esophagus when screening with forceps biopsies. Esophagus 2023;20:143-9. [Crossref] [PubMed]
- Yalamarthi S, Witherspoon P, McCole D, et al. Missed diagnoses in patients with upper gastrointestinal cancers. Endoscopy 2004;36:874-9. [Crossref] [PubMed]
- Ha VD, Quach TD. Rate and characteristics of interval gastric carcinoma with advanced endoscopic features. Ho Chi Minh Journal of Medicine. 2018;22:56-62.
- Arif AA, Jiang SX, Byrne MF. Artificial intelligence in endoscopy: Overview, applications, and future directions. Saudi J Gastroenterol 2023;29:269-77. [Crossref] [PubMed]
- Vu H, Manh XH, Duc BQ, et al. editors. Labelling stomach anatomical locations in upper gastrointestinal endoscopic images using a CNN. Proceedings of the 10th International Symposium on Information and Communication Technology; 2019.
- Manh ND, Hang DV, Long DV, et al. editors. EndoUNet: A Unified Model for Anatomical Site Classification, Lesion Categorization and Segmentation for Upper Gastrointestinal Endoscopy. 2022 14th International Conference on Knowledge and Systems Engineering (KSE); 2022.
- Tran TH, Nguyen PT, Tran DH, et al. editors. Classification of anatomical landmarks from upper gastrointestinal endoscopic images⋆. 2021 8th NAFOSTED Conference on Information and Computer Science (NICS); 2021.
- Vu H, Nguyen Xuan C, Le MN, et al. editors. A robust and high-performance neural network for classifying landmarks in upper gastrointestinal endoscopy images. Proceedings of the 11th International Symposium on Information and Communication Technology; 2022.
- Romańczyk M, Ostrowski B, Lesińska M, et al. The prospective validation of a scoring system to assess mucosal cleanliness during EGD. Gastrointest Endosc 2024;100:27-35. [Crossref] [PubMed]
- Tan M, Le Q. editors. EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning; 2019.
.Tan M Le QV EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv. [Preprint].2020 . doi: .- Deng J, Dong W, Socher R, et al. editors. ImageNet: A large-scale hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern Recognition; 20-25 June 2009; Miami, FL, USA. IEEE; 2009.
Reis D Kupec J Hong J Real-Time Flying Object Detection with YOLOv8. arXiv. [Preprint].2024 . doi: .Lin TY Maire M Belongie S Microsoft COCO: Common Objects in Context. arXiv. [Preprint].2015 . doi: .- Mülder DT, Hahn AI, Huang RJ, et al. Prevalence of Gastric Precursor Lesions in Countries With Differential Gastric Cancer Burden: A Systematic Review and Meta-analysis. Clin Gastroenterol Hepatol 2024;22:1605-1617.e46. [Crossref] [PubMed]
- Yin Y, Liang H, Wei N, et al. Prevalence of chronic atrophic gastritis worldwide from 2010 to 2020: an updated systematic review and meta-analysis. Ann Palliat Med 2022;11:3697-703. [Crossref] [PubMed]
- Dao Viet Hang, Thi Minh Hue Luu, Duy Thang Nguyen. The Prevalence of esophageal motility disorders based on Chicago 4.0 classification in patients with upper gastrointestinal symptoms. Vietnam Medical Journal 2024;543: [Crossref]
- Dent J, Becher A, Sung J, et al. Systematic review: patterns of reflux-induced symptoms and esophageal endoscopic findings in large-scale surveys. Clin Gastroenterol Hepatol 2012;10:863-873.e3. [Crossref] [PubMed]
- Groenen MJ, Kuipers EJ, Hansen BE, et al. Incidence of duodenal ulcers and gastric ulcers in a Western population: back to where it started. Can J Gastroenterol 2009;23:604-8. [Crossref] [PubMed]
- Romańczyk M, Ostrowski B, Kozłowska-Petriczko K, et al. Scoring system assessing mucosal visibility of upper gastrointestinal tract: The POLPREP scale. J Gastroenterol Hepatol 2022;37:164-8. [Crossref] [PubMed]
- Bisschops R, Areia M, Coron E, et al. Performance measures for upper gastrointestinal endoscopy: a European Society of Gastrointestinal Endoscopy (ESGE) Quality Improvement Initiative. Endoscopy 2016;48:843-64. [Crossref] [PubMed]
- Armstrong D. Endoscopic evaluation of gastro-esophageal reflux disease. Yale J Biol Med 1999;72:93-100.
- Manabe N, Joh T, Higuchi K, et al. Clinical significance of gastroesophageal reflux disease with minimal change: a multicenter prospective observational study. Sci Rep 2022;12:15036. [Crossref] [PubMed]
- Dixon MF, Genta RM, Yardley JH, et al. Classification and grading of gastritis. The updated Sydney System. International Workshop on the Histopathology of Gastritis, Houston 1994. Am J Surg Pathol 1996;20:1161-81. [Crossref] [PubMed]
- Toyoshima O, Nishizawa T, Koike K. Endoscopic Kyoto classification of Helicobacter pylori infection and gastric cancer risk diagnosis. World J Gastroenterol 2020;26:466-77. [Crossref] [PubMed]
- Ohya TR, Endo H, Kawagoe K, et al. A prospective randomized trial of lafutidine vs rabeprazole on post-ESD gastric ulcers. World J Gastrointest Endosc 2010;2:36-40. [Crossref] [PubMed]
- Japanese Classification of Esophageal Cancer, 11th Edition: part I. Esophagus 2017;14:1-36.
- Japanese classification of gastric carcinoma: 3rd English edition. Gastric Cancer 2011;14:101-12. [Crossref] [PubMed]
- Nagtegaal ID, Odze RD, Klimstra D, et al. The 2019 WHO classification of tumours of the digestive system. Histopathology 2020;76:182-8. [Crossref] [PubMed]
- Bisschops R, Areia M, Coron E, et al. Performance measures for upper gastrointestinal endoscopy: A European Society of Gastrointestinal Endoscopy quality improvement initiative. United European Gastroenterol J 2016;4:629-56. [Crossref] [PubMed]
- Fernández-Esparrach G, Marín-Gabriel JC, Díez Redondo P, et al. Quality in diagnostic upper gastrointestinal endoscopy for the detection and surveillance of gastric cancer precursor lesions: Position paper of AEG, SEED and SEAP. Gastroenterol Hepatol 2021;44:448-64. [Crossref] [PubMed]
- Huang Q, Shi J, Sun Q, et al. Clinicopathological characterisation of small (2 cm or less) proximal and distal gastric carcinomas in a Chinese population. Pathology 2015;47:526-32. [Crossref] [PubMed]
- Choi SJ, Khan MA, Choi HS, et al. Development of artificial intelligence system for quality control of photo documentation in esophagogastroduodenoscopy. Surg Endosc 2022;36:57-65. [Crossref] [PubMed]
- Wu L, Zhang J, Zhou W, et al. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy. Gut 2019;68:2161-9. [Crossref] [PubMed]
- Ge H, Zhou X, Wang Y, et al. Development and Validation of Deep Learning Models for the Multiclassification of Reflux Esophagitis Based on the Los Angeles Classification. J Healthc Eng 2023;2023:7023731. [Crossref] [PubMed]
- Gulati S, Bernth J, Liao J, et al. OTU-07 Near focus narrow and imaging driven artificial intelligence for the diagnosis of gastro-oesophageal reflux disease. Gut 2019;68:A4.
- Wen Y, Huang Y, Liu Y, et al. Artificial intelligence for the diagnosis of Helicobacter pylori infection in endoscopic and pathological tissues images: A systematic review and meta-analysis. Intelligence-Based Medicine 2025;11:100244.
- Turtoi DC, Brata VD, Incze V, et al. Artificial Intelligence for the Automatic Diagnosis of Gastritis: A Systematic Review. J Clin Med 2024;13:4818. [Crossref] [PubMed]
- Shi Y, Wei N, Wang K, et al. Diagnostic value of artificial intelligence-assisted endoscopy for chronic atrophic gastritis: a systematic review and meta-analysis. Front Med (Lausanne) 2023;10:1134980. [Crossref] [PubMed]
- Parkash O, Lal A, Subash T, et al. Use of artificial intelligence for the detection of Helicobacter pylori infection from upper gastrointestinal endoscopy images: an updated systematic review and meta-analysis. Ann Gastroenterol 2024;37:665-73. [Crossref] [PubMed]
- Wu L, He X, Liu M, et al. Evaluation of the effects of an artificial intelligence system on endoscopy quality and preliminary testing of its performance in detecting early gastric cancer: a randomized controlled trial. Endoscopy 2021;53:1199-207. [Crossref] [PubMed]
- Sumiyama K, Futakuchi T, Kamba S, et al. Artificial intelligence in endoscopy: Present and future perspectives. Dig Endosc 2021;33:218-30. [Crossref] [PubMed]
- de Groof AJ, Struyvenberg MR, Fockens KN, et al. Deep learning algorithm detection of Barrett's neoplasia with high accuracy during live endoscopic procedures: a pilot study (with video). Gastrointest Endosc 2020;91:1242-50. [Crossref] [PubMed]
- Mori Y, Kudo SE, Mohmed HEN, et al. Artificial intelligence and upper gastrointestinal endoscopy: Current status and future perspective. Dig Endosc 2019;31:378-88. [Crossref] [PubMed]
- Leenhardt R, Fernandez-Urien Sainz I, Rondonotti E, et al. PEACE: Perception and Expectations toward Artificial Intelligence in Capsule Endoscopy. J Clin Med 2021;10:5708. [Crossref] [PubMed]
- Wadhwa V, Alagappan M, Gonzalez A, et al. Physician sentiment toward artificial intelligence (AI) in colonoscopic practice: a survey of US gastroenterologists. Endosc Int Open 2020;8:E1379-84. [Crossref] [PubMed]
- Messmann H, Bisschops R, Antonelli G, et al. Expected value of artificial intelligence in gastrointestinal endoscopy: European Society of Gastrointestinal Endoscopy (ESGE) Position Statement. Endoscopy 2022;54:1211-31. [Crossref] [PubMed]
- Ahmed Z, Bhinder KK, Tariq A, et al. Knowledge, attitude, and practice of artificial intelligence among doctors and medical students in Pakistan: A cross-sectional online survey. Ann Med Surg (Lond) 2022;76:103493. [Crossref] [PubMed]
Cite this article as: Dao HV, Nguyen HN, Duong GQ, Tran DA, Nguyen HV, Dao KT, Tran MN, Nguyen BP, Nguyen TT, Lam HN, Nguyen TTH, Hoang LB, Truong SQH, Dao LV. A real-time artificial intelligence-integrated system (DrAidTM Endo) feasibility in identifying anatomical landmarks and detecting upper gastrointestinal tract lesions: a randomized controlled trial. J Med Artif Intell 2026;9:42.


