The feasibility of using artificial intelligence-powered translation software in lieu of human interpreters for low-risk, Arabic-speaking pregnant people in an Australian antenatal clinic
Original Article

The feasibility of using artificial intelligence-powered translation software in lieu of human interpreters for low-risk, Arabic-speaking pregnant people in an Australian antenatal clinic

Josephine S. F. Chow1,2,3,4,5, Nutan Maurya1,2, Sam Shen6, Shivani Mani1,2, Sharon May7, Brenda Gillard8, Jon Hyett9, Carlos Bowkett10, Lewis Bennett10, Aleisha Heys7

1South Western Sydney Nursing & Midwifery Research Alliance (SWS NMRA), South Western Sydney Local Health District, Liverpool, Australia; 2Robotics and Innovation in Bio-sensing and Technology (ribT), Ingham Institute for Applied Medical Research, Liverpool, Australia; 3Faculty of Medicine, University of New South Wales, Sydney, Australia; 4NICM, Western Sydney University, Sydney, Australia; 5Faculty of Health Science, University of Tasmania, Hobart, Australia; 6Multicultural Services, South Western Sydney Local Health District, Liverpool, Australia; 7Antenatal Clinics, Fairfield Hospital, South Western Sydney Local Health District, Liverpool, Australia; 8Clinical Governance Unit, South Western Sydney Local Health District, Liverpool, Australia; 9Department of Obstetrics and Gynaecology, Faculty of Health, Western Sydney University, Sydney, Australia; 10Office of the Chief Scientist & Engineer, Premier’s Department, NSW Government, Sydney, Australia

Contributions: (I) Conception and design: JSF Chow, B Gillard, J Hyett, C Bowkett, L Bennett; (II) Administrative support: N Maurya, S Mani; (III) Provision of study materials or patients: JSF Chow, A Heys, S Shen, N Maurya, S Mani; (IV) Collection and assembly of data: N Maurya; (V) Data analysis and interpretation: JSF Chow, N Maurya; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Josephine S. F. Chow, PhD, MBA. South Western Sydney Nursing & Midwifery Research Alliance (SWS NMRA), South Western Sydney Local Health District, Liverpool, Australia; Robotics and Innovation in Bio-sensing and Technology (ribT), Ingham Institute for Applied Medical Research, 1 Campbell Street, Liverpool, NSW 2170, Australia; Faculty of Medicine, University of New South Wales, Sydney, Australia; NICM, Western Sydney University, Sydney, Australia; Faculty of Health Science, University of Tasmania, Hobart, Australia. Email: Josephine.Chow@health.nsw.gov.au.

Background: Language barriers in maternity care can compromise patient safety and experience, particularly for culturally and linguistically diverse (CALD) populations. Mobile transcription and translation technologies have emerged as potential solutions for low-risk clinical interactions; however, evidence on their real-world usability in antenatal settings remains limited. This study aimed to evaluate the feasibility and usability of a real-time transcription and translation software in facilitating communication between clinicians and Arabic-speaking pregnant women requiring interpreter services.

Methods: A feasibility study was conducted in the antenatal clinic at a teaching hospital. Ten Arabic-speaking women attending low-risk antenatal appointments were recruited. Each participant completed two consultations: the first using standard interpreter services and the second using the software with interpreter oversight. Pre- and post- study surveys captured participant experience, while interpreters assessed translation accuracy using a structured tool adapted from New South Wales (NSW) Multicultural Health Communication Service guidelines. Data were analysed descriptively.

Results: Half of the participants reported being satisfied/very satisfied with software-mediated communication, and 60% experienced no difficulty understanding translations. However, interpreter assessments indicated poor translation accuracy, with mean clarity and meaning scores of 3.3 and 3.2 (scale 1–10), and factual errors in 60% of cases. Connectivity emerged as a critical factor, with network disruptions causing delays and impacting usability.

Conclusions: Findings demonstrate the feasibility of the software in real-world, low-risk antenatal consultations and its ability to support two-way communication. However, concerns regarding translation accuracy, reliability, and connectivity currently limit its clinical utility, necessitating ongoing interpreter supervision and substantial refinement before broader implementation. Notably, the discrepancy between participant satisfaction and interpreter-identified errors, together with reliance on interpreter assessments, introduces the potential for observer, expectation, and vested interest biases, and therefore warrants cautious interpretation of the findings.

Keywords: Translation; transcription; artificial intelligence (AI); antenatal consultations


Received: 06 March 2026; Accepted: 18 June 2026; Published online: 27 July 2026.

doi: 10.21037/jmai-2026-0054


Highlight box

Key findings

• This study demonstrated that artificial intelligence (AI)-based translation software is feasible for supporting two-way communication in low-risk antenatal consultations. While 50% of participants reported satisfaction and 60% experienced no difficulty understanding translations, interpreter assessments identified poor translation accuracy, including factual errors and clinically significant discrepancies. Connectivity issues further affected usability. The discrepancy between patient-reported experience and interpreter-evaluated accuracy, alongside reliance on interpreter assessment, introduces potential observer, expectation, and vested-interest biases. Overall, these findings indicate that the software is not yet suitable for independent clinical use and requires further refinement.

What is known and what is new?

• Mobile translation applications are increasingly utilized in healthcare to mitigate language barriers, when professional interpreters are unavailable. However, evidence regarding their effectiveness in antenatal care remains scarce.

• This study provides novel real-world evidence evaluating AI-driven, real-time transcription and translation software in an antenatal setting. It demonstrates feasibility in facilitating two-way communication in controlled, low-risk scenarios with appropriate clinical oversight, particularly where interpreter services are unavailable or delayed. Nevertheless, translation inaccuracies and technical limitations remain important concerns.

What is the implication, and what should change now?

• The software may serve as a supplementary communication tool in low-risk, well-defined clinical scenarios, with clinician oversight. Healthcare services should establish clear governance frameworks, including standard operating procedures, staff training, and escalation pathways to interpreter services. Ongoing system improvement, validation, and monitoring are required to ensure safe, reliable, and equitable clinical use. Further research should prioritise improving translation accuracy, system reliability, and connectivity, before considering broader implementation.


Introduction

Background

In Australia, people who are born overseas and/or speak a language other than English at home are classified as culturally and linguistically diverse (CALD). In 2021, more than 7 million Australians were born overseas, representing 27.6% of the population, and 5.8 million (22.8%) reported using a language other than English at home (1). The proportion of overseas-born residents has steadily increased across all states and territories. In New South Wales (NSW), 29.3% of the population was born overseas, and over 27% reported speaking a language other than English at home (1,2).

People from CALD backgrounds often experience greater challenges navigating the healthcare system (3-5). NSW Health is the public healthcare system of the Australian state of NSW. It comprises a network of local health districts, specialty health networks, and supporting agencies, and is responsible for delivering a broad range of services including hospital care, community health, and population health programs. NSW Health aims to provide safe, high-quality, and compassionate healthcare to the population.

NSW Health has committed to improving healthcare access for CALD communities through initiatives such as the NSW Plan for Healthy CALD Communities: 2019–2023 (6). Current strategies include multicultural health communication services and dedicated language service contracts for each local health district (2,7). However, districts with high proportions of English-as-a-second language (ESL) speakers face significant pressure to provide adequate services. South Western Sydney Local Health District (SWSLHD), one of Australia’s most culturally diverse regions, exemplifies these challenges, particularly in delivering cost-efficient services to CALD patients (8,9). Demand for medical interpretation services is high, and the diversity of languages poses difficulties for nurses and midwives, especially in non-acute settings (10,11).

Mobile translation applications (MTAs) have emerged as a potential solution for facilitating timely communication in low-risk scenarios when professional interpreters are unavailable (10,12-15). Several MTAs have been trialed in healthcare settings, including CALD Assist, Talk to Me, Google Translate, Microsoft Translator, Travis Translator, and Apple iTranslate (12,13,16-19). CALD Assist and Talk to Me, developed in Australia, are tailored for healthcare settings and use fixed phrases with images and audio to support communication with patients with limited English proficiency (LEP) (10,12,16-18). These applications are effective for simple interactions but lack flexibility for complex or sensitive conversations. Conversely, real-time voice-to-voice applications such as Google Translate and Microsoft Translator are widely accessible but yield mixed results in clinical contexts. While some studies report user satisfaction (15,19,20), others highlight poor accuracy, particularly for medical terminology and certain languages (21-23). These tools perform best with short, simple phrases, limiting their suitability to low-risk settings (24).

Rationale and knowledge gap

Translation applications have been used to deliver pregnancy-related education and culturally sensitive information during antenatal care, improving access and patient satisfaction (25-27). However, while such technologies exist, there is a lack of published evidence evaluating real-time voice-to-voice translation tools in clinical antenatal care settings.

Artificial intelligence (AI)-driven initiatives in Australia include multilingual video translation, hand-held devices, and MTAs such as CALD Assist and Talk to Me (12-14,16,17). While these tools support low-risk communication, evidence on their real-world effectiveness, particularly in antenatal care, remains limited. There is an urgent need for technological solutions that deliver accurate, responsive, and equitable health information for CALD communities (14,15,28).

Transcription & translation software

To address the issue of interpreter shortages and communication barriers experienced by CALD population in South Western Sydney (SWS), in 2022, lead author (J.S.F.C.), was awarded funding under the NSW Small Business Innovation and Research (SBIR) program, led by the Office of the Chief Scientist and Engineer, to explore AI-enabled solutions that support health service delivery for CALD communities. The Secretary of NSW Health and the former Chief Executive of SWSLHD endorsed SWSLHD’s participation in the program. The first phase involved co-design and proof-of-concept studies conducted with selected innovator in 2022, progressing to feasibility study in 2024.

A transcription and translation software was subsequently developed to support communication between clinicians and CALD patients. This software provides voice-to-voice and voice-to-text translation functionality. It was co-designed with midwives, pregnant women with live-experienced, health service language staff and the research team. The software integrates real-time transcription, translation, and voice generation functionalities, powered by an ensemble of AI models with customizable vocabularies. It is designed to facilitate two-way communication between clinicians and CALD patients and provide a user-friendly interface for administrating common patient-reported outcome measures.

Objective

This feasibility study aimed to evaluate the feasibility and usability of an AI-enabled transcription and translation software to support communication between clinicians, interpreters and patients who speak a language other than English within the antenatal care context. We present this article in accordance with the STROBE reporting checklist (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0054/rc).


Methods

Study design & setting

This feasibility study was designed as an exploratory, small-scale investigation to determine whether the software is feasible to implement and usable in a real-world clinical setting prior to larger study.

This study was conducted in the antenatal clinic at Fairfield Hospital in Australia. The study assessed the performance of the transcription and translation software during routine antenatal consultations.

Data source and population

Given the demographics of the study site and acknowledging that Arabic includes multiple dialects which may impose implementation challenge, Arabic-speaking pregnant women attending low-risk antenatal appointments and requiring interpreter services were recruited between June and July 2025.

The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by SWSLHD Human Ethics Research Committee (2024/ETH01838). Written informed consent was obtained from all participants prior to any study procedure.

Low-risk antenatal care for this study was defined in accordance with the Australian College of Midwives (ACM) Guidelines for Consultation and Referral (29). Women were eligible if they were receiving care under a midwifery-led model, classified as Category A, in which the midwife acts as the primary care provider and coordinates all aspects of care. Consultation with a midwifery colleague, medical practitioner, or other healthcare providers may occur if clinically indicated.

A pregnancy is generally considered low risk when the individual is in good overall health, with no significant pre-existing medical conditions (e.g., diabetes, hypertension, or cardiac disease), and where any previous pregnancies have been uncomplicated. Additional characteristics of low-risk pregnancy include the absence of concerns regarding fetal growth or development, a singleton pregnancy, and no significant complications arising during the current pregnancy, such as preeclampsia, gestational diabetes, or preterm labor.

All women included in this study were assessed by the midwife as suitable for ongoing low-risk antenatal care at the time of consultation, with no immediate complications requiring escalation to specialist or high-risk obstetric services.

Sample size

As a feasibility study, hypothesis testing and statistical power calculations were not applicable. A target sample size of 10 participants was set to evaluate the feasibility and usability of the solution.

Study procedures

Participation was voluntary. Screening for eligibility was conducted by a clinical midwifery consultant or clinical midwifery educator based on predefined inclusion and exclusion criteria (Figure 1).

Figure 1 Flowchart representing study methodology. CF, consent form; PIS, participant information sheet.

Inclusion criteria:

  • Pregnant women attending antenatal clinic;
  • Low-risk antenatal care as determined by clinical assessment;
  • Requirement for interpreter services during consultation;
  • Willingness to provide written informed consent and comply with study procedures.

Exclusion criteria:

  • High-risk antenatal care as determined by clinical assessment;
  • Decline to provide written informed consent.

Eligible participants received a participant information sheet (PIS) and consent form (CF). An interpreter from Health Language Services, accredited by the National Accreditation Authority for Translators and Interpreters (NAATI), was present during recruitment and consultations to explain study details and address participant queries. All study documents, including PIS, CF and survey questions, were translated into Arabic by accredited NAATI staff.

Consultation process

Each participant attended two consultations.

First consultation—standard care

Participants were screened for eligibility and provided with study information. Interpreter explained the study and answered questions. Written consent was obtained prior to enrolment.

Second consultation—intervention

The AI translation software was used in real-time, to facilitate two-way conversation between the clinician and participant. An interpreter was present throughout the consultation to monitor any critical issues and to assess translation accuracy using a structured evaluation tool.

During consultations, participants experienced only voice-to-voice translation. Audio from the consultation was automatically transcribed by the software, generating text-based transcripts. These transcripts were not accessible to the participants during the consultation and were used solely for evaluation purposes. AI-generated transcripts were reviewed by the interpreter post-consultation for accuracy assessment using the structured evaluation tool. The interpreter’s assessment was based on the real-time observation of the interaction and AI-generated transcription output.

The participants consultations were not identical in content, as they reflected routine care aligned with the participant’s gestational age and individual clinical needs. The timing and interval between the two consultations were also not fixed, as appointments were scheduled based on routine antenatal care pathways and the gestational age of each participant. Midwives organised consultations in accordance with standard clinical scheduling, and researchers attended the antenatal clinic on the day of the participant’s scheduled appointment to conduct the intervention.

The antenatal consultations were routine, planned interactions with a midwife and involved low-complexity clinical and informational discussions. The primary purposes of these visits included monitoring maternal and foetal wellbeing, early identification of emerging complications, provision of health education, and coordination of ongoing care.

Typical consultation activities included:

  • Blood pressure measurement;
  • Weight (at key time points, not necessarily every visit);
  • Review of symptoms (e.g., bleeding, pain, reduced fetal movements later in pregnancy);
  • Review of results from previous investigations;
  • Foetal heart rate auscultation and movements;
  • Assessment of foetal movements (after quickening);
  • Fundal height measurement (from ~24 weeks);
  • Abdominal palpation;
  • Ongoing risk assessment to confirm continued suitability for low-risk pathway;
  • Identification of any deviation requiring escalation/referral;
  • Review of medical, obstetric, and psychosocial factors;
  • Review of ultrasound;
  • Screening for gestational diabetes (~24–28 weeks).

No consultations involved emergency care, high-risk clinical scenarios, or complex decision-making processes, such as diagnosis of complications, urgent management decisions, or high-stakes informed consent discussions.

Participant surveys

To capture participant feedback and inform future improvement to the software, a pre- and post- study survey were administered.

Pre-study survey

Prior to the intervention, participants completed a brief survey to document their previous experience using any transcription or translation software for communication with the health care providers (Table 1).

Table 1

Pre-study participant survey

No. Survey questions   Response options
1 Have you ever used any transcription and translation software to communicate with your health care provider?   Yes/No
1.1 If yes: did you find the transcription and translation software useful to communicate with your health care provider?   Very useful/somewhat useful/not useful
1.2 Did you have any difficulty in understanding or communicating with your health care provider via the transcription and translation software?   Yes/no, if yes, please specify
1.3 How was your overall experience communicating via the transcription and translation software?   Extremely unhappy/dissatisfied/neutral/satisfied/extremely satisfied
2 Given an option, how would you prefer to communicate with your health care provider?   Via human interpreter/via transcription and translation software
Post-study survey

Following the second consultation, participants completed a post-study survey evaluating their experience with the transcription and translation software during the consultation. The survey included both closed-ended questions (using rating scales and yes/no responses) and open-ended questions to gather qualitative insights.

The post-study survey assessed multiple domains related to participants’ experience with the AI transcription and translation software (Table 2). These included overall satisfaction with communication via the software, clarity and ease of understanding of translated content, and any difficulties compared to communication through a human interpreter. Participants were also asked to identify perceived differences between software-mediated and interpreter-mediated interactions, indicate their preferred communication method for future consultations, and rate the software’s overall performance. Finally, open-ended questions captured suggestions for improvement to enhance usability and effectiveness.

Table 2

Post-study participant survey

No. Survey questions   Response options
1 How was your experience communicating to your health care provider via the transcription & translation software?   Very satisfied/satisfied/neither; dissatisfied/very dissatisfied
2 Was the translation clear and easy to understand?   Yes/no, if no, please specify
3 Did you have any difficulty/problem understanding the language translated by the software as compared to human Interpreter?   Yes/no, if yes, please specify
4 What difference did you notice when communicating via the transcription & translation software compared to human Interpreter?
5 Which means/way do you think was easier and more convenient to communicate with your health care provider?   Via human Interpreter/via Software
6 How would you rate this transcription & translation software?   Outstanding/exceeds expectations/meets expectations/needs improvement/unacceptable
7 If you have to suggest any improvement, what would it be?
8 Given an option, which means/way would you prefer in future to communicate with your health care provider?   Via human interpreter/via software

Interpreter assessment

Two interpreters were involved in the study, with cases distributed between them based on their availability rather than having both interpreters assess all cases. To ensure consistency, both interpreters followed the same structured assessment framework and were provided with standardised guidance prior to data collection. The interpreter evaluated the accuracy of the software during the consultation and reviewed the transcription post-consultation using a structured assessment tool. The questions in the tool were adapted from Translation Checking Guidelines developed by the NSW Multicultural Health Communication Service (MHCS) (30) and modified to align with the study objectives. The assessment focused on four key criteria: (I) whether the translation was clear and easily understood; (II) whether the translation conveyed the same meaning as the original English text; (III) whether the translation contained any factual errors; and (IV) whether any information was misleading to the extent that it could cause harm.

Data security and privacy

Data security and privacy were managed in accordance with institutional and regulatory requirements. The transcription and translation software was developed and provided by an independent vendor and operated via a secure cloud-based platform. Data transmitted during consultations were protected using encrypted communication protocols to minimize the risk of unauthorized access.

No identifiable patient data was intentionally retained within the system beyond what was required for real-time processing. AI-generated transcripts used for evaluation were de-identified prior to analysis and stored securely on password-protected systems within hospital network drive, accessible only to authorized study personnel.

The use of the vendor platform was governed by local data governance and ethics approvals, ensuring alignment with NSW Health policies and the Australian Privacy Principles (APPs). The vendor did not have access to identifiable study data for purposes outside the scope of service delivery.

Participants were informed of the use of the software and the handling of their information as part of the consent process.

Statistical analyses

Given the exploratory nature of this study, all statistical analyses were descriptive rather than inferential. Survey responses were entered into Microsoft Excel and analyzed using descriptive statistics, including frequencies, percentages, and measures of central tendency (where applicable) to summarize participant feedback. Closed-ended questions were reported as categorical data, while open-ended responses from the post-study survey were analyzed using thematic analysis. Responses were reviewed and coded inductively to identify recurring ideas and patterns. Similar responses were grouped into categories, and common themes were summarized descriptively to complement quantitative findings. Due to the small sample size, the analysis was descriptive in nature and aimed to highlight common patterns in participant feedback rather than achieve theoretical saturation.

Interpreter assessment data were similarly summarized using descriptive statistics to evaluate translation clarity, accuracy, and potential safety concerns. No hypothesis testing or formal power calculations were performed, as the primary aim was to assess feasibility and usability.


Results

A total of 10 Arabic-speaking pregnant women attending antenatal clinic at the participating hospital consented and were recruited for the study. The mean age was 30 years (range, 22–37 years), and 50% (n=5) of the participants had one or more children.

Survey data

Pre-study participant survey

Pre-study survey data demonstrated that none of the participants had previous experience of using any transcription and translation software.

Post-study participant survey

Post-study survey data showed that 50% (n=5) of the participants were very satisfied or satisfied with communication during via software the consultation and indicated that the translation was clear and easy to understand. Sixty percent (n=6) reported no difficulty understanding the language translated by the software compared to a human interpreter. When asked about differences between software-mediated and interpreter-mediated communication, 50% (n=5) stated they noticed no difference.

In terms of performance ratings, 20% (n=2) of the participants rated the software as “Outstanding”, 10% (n=1) rated it as “Meets Expectations”, and 70% (n=7) indicated it “Needs Improvement”. Suggestions for improvement included clearer instructions on when to start or stop conversations, use of more suitable and clear words, stronger network connectivity, and reducing language-time lag. When asked about future use, 30% (n=3) of participants indicated they would prefer to use the software again, while the remainder favoured human interpreters (Table 3).

Table 3

Post-study participant survey results (closed-ended questions)

No. Survey questions   Response options
1 How was your experience communicating to your health care provider via the transcription & translation software?   Very satisfied/satisfied: 5; neither: 2; dissatisfied: 3
2 Was the translation clear and easy to understand?   Yes: 5; no: 5
3 Did you have any difficulty/problem understanding the language translated by the software as compared to human Interpreter?   Yes: 4; no: 6
4 Which means/way do you think was easier and more convenient to communicate with your health care provider?   via human Interpreter: 8; via Software: 2
5 How would you rate this transcription & translation software?   Outstanding: 2; meets expectations: 1; needs improvement: 7
6 Given an option, which means/way would you prefer in future to communicate with your health care provider?   via human Interpreter: 7; via Software: 3

Translation & transcription software usage & challenges

The average duration of software use during consultations was 10 minutes (range, 5–20 minutes). Participant experience and satisfaction varied, influenced by several factors:

  • Internet connectivity—initial delays in translation were attributed to poor network connectivity. Switching from the hospital’s guest Wi-Fi to a 5G data subscriber identity module (SIM) and whitelisting the device improved performance. However, occasional disruptions requiring link resets caused delays and interrupted communication flow.
  • Translation accuracy—accuracy was affected by patient speech characteristics, including soft voices, use of dialects, and mixing English with Arabic words. Ambiguity due to words with multiple meanings also impacted translation quality.
  • User familiarity—lack of familiarity with the software led to errors and slower communication. For example, participants were sometimes unsure when to start speaking, resulting in overlapping voices and translation inaccuracies.
  • Environmental factors—external interruptions, such as family members interjecting during consultations, disrupted conversational flow and contributed to misinterpretation.

Interpreter assessment data

The interpreters (n=2) assessment revealed limitations in the software’s translation accuracy and clarity (Table 4). Only one instance was rated as clear and easily understood, and only one translation was judged to convey the same meaning as the English original. Mean scores for clarity (3.3/10) and meaning alignment (3.2/10) indicate poor overall performance. Furthermore, factual errors were identified in 60% (n=6) of cases, and 40% (n=4) of translations were considered potentially misleading to the extent of causing harm. These findings highlight safety concerns and underscore the risk of relying solely on automated translation tools in clinical settings without professional oversight. The unanimous recommendation for changes (100%) reinforces the need for substantial improvements in translation accuracy, user guidance, and system reliability before broader implementation.

Table 4

Outcome from interpreter assessment on the translations

No. Survey questions Response options
1 Is the translation clear and easily understood? No: 9; yes: 1
2 On a scale of 1–10, please rate your overall understanding of the translation: mean (SD), range (min–max) 3.3 (1.64), 2–6
3 Does the translation communicate the same meaning as the English original? No: 9; yes: 1
4 On a scale of 1–10, please rate the translation communication comparing in English: mean (SD), range (min–max) 3.2 (1.62), 2–6
5 Does the translation contain factual errors? No: 4; yes: 6
6 Is any information so misleading it might cause harm? No: 6; yes: 4
7 Are changes recommended for this resource? Yes: 10

max, maximum; min, minimum; SD, standard deviation.


Discussion

Key findings

This study evaluated the feasibility and usability of a co-designed AI transcription and translation software in an antenatal setting. Overall, 50% of participants reported being satisfied or very satisfied with their experience using the software, and 60% indicated no difficulty understanding the translated language compared to a human interpreter. The experience and satisfaction of the participants were influenced by factors such as internet connectivity, translation accuracy, user familiarity, and environmental conditions. Half of the participants perceived no difference between software-mediated and interpreter-mediated communication, and approximately one-third expressed willingness to use the software in future consultations, provided improvements were made. Suggested enhancements included clearer instructions for use, improved word selection, stronger network connectivity, and reduced time lag.

A key finding was the discrepancy between participant-reported satisfaction and interpreter-rated accuracy. While participants perceived communication as adequate, professional assessment identified substantial inaccuracies and potential safety risks. This highlights a critical distinction between perceived understanding and actual comprehension in clinical communication.

Taken together, the results suggest that while the software may enable basic communication, its current performance is insufficient for safe, independent clinical use.

Strengths and limitations

This study addresses an important gap by evaluating real-time transcription and translation software within a real-world antenatal care setting, an area where evidence remains limited. The feasibility design allowed for direct observation of real-world usability during antenatal consultations, providing practical insights into implementation challenges and user experience. Inclusion of both patient feedback and interpreter assessment strengthened the evaluation by incorporating multiple perspectives on translation accuracy and safety. Additionally, the study adhered to rigorous ethical standards and used validated assessment criteria adapted from NSW MHCS guidelines, enhancing the reliability of findings.

However, several limitations should be acknowledged. The study involved a small sample size (n=10) and was conducted at a single site, limiting generalisability. The evaluation was restricted to Arabic-speaking participants, and variation in dialects may have influenced translation accuracy. Another limitation is the absence of clinician feedback, which may influence interpretation of the findings. Clinicians are primary end-users of translation technologies and play a central role in mediating communication, assessing patient understanding, and making clinical decisions. Without their perspectives, it is not possible to fully evaluate the software’s impact on clinical workflow integration, communication efficiency, or decision-making processes. Future studies should incorporate structured clinician evaluation alongside patient and interpreter assessments.

A further limitation relates to the lack of blinding in interpreter assessments, which may have introduced observer and expectation bias. Interpreters were aware that the translations were AI generated, which may have influenced their evaluations. It is possible that professional interpreters, consciously or unconsciously, applied more stringent criteria or were more attuned to identifying errors when assessing AI generated translations. Additionally, given that translation technologies may be perceived as a potential challenge to professional interpreting roles, there is a theoretical possibility of vested interest bias, whereby evaluators may emphasise the limitations of automated tools.

While interpreters are trained to adhere to objective standards of accuracy and fidelity, these potential influences should be considered when interpreting the findings. Importantly, the low rating scores may reflect both genuine deficiencies in translation quality and stricter professional benchmarks applied during evaluation. Furthermore, most assessments (90%) were conducted by a single interpreter, limiting assessment reliability. Future studies should aim to mitigate this risk by incorporating blinded assessment procedures and involving multiple independent raters to improve objectivity and inter-rater reliability.

Given the small sample size and exploratory design, findings should be interpreted cautiously and are not intended to establish effectiveness or generalisability.

Comparison with similar research

These findings align with previous studies that have highlighted the potential of MTAs to address language barriers in health care, particularly for low-risk conversations when professional interpreters are unavailable (10,12,15-17,31). However, consistent with earlier research, usability challenges remain, including translation inaccuracies, risk of misinterpretation in high-risk scenarios, and issues related to connectivity and user familiarity (18,22,32). In our study, factors such as internet connectivity, patient accent or dialect, and lack of familiarity with the software likely influenced user experience and satisfaction. This study contributes novel evidence within an antenatal care context, highlighting both the potential utility and significant safety limitations of AI-mediated translation in real-world clinical environment.

Explanations of findings

This feasibility study was conducted to determine whether AI-powered transcription and translation software can be feasibly implemented in practice and demonstrate early indications of usability and potential value, rather than establish clinical effectiveness. The findings demonstrated feasibility of software in facilitating two-way communication during antenatal consultations, and a proportion of participants reported satisfactory user experience. However, limitations were identified, particularly in translation accuracy, reliability, and dependence on connectivity, with interpreter assessments highlighting clinically significant errors. Taken together, these results suggested that while the concept of real-time AI-mediated translation in antenatal care is operationally feasible and acceptable to some users, it is not yet sufficiently reliable or accurate to be considered clinically safe for independent use. This study supports the feasibility and usability components, but does not support readiness for broader clinical implementation, reinforcing the need for further refinement and rigorous evaluation.

Internet connectivity emerged as a critical factor influencing performance, particularly during initial trials using the organization’s guest Wi-Fi, which resulted in delays and disrupted conversational flow. Switching to a 5G data SIM and whitelisting the device improved speed and stability; however, occasional interruptions still required reconnecting, causing further delays. These connectivity challenges highlight that real-time cloud-based translation technologies are highly sensitive to bandwidth and network reliability, making infrastructure readiness a critical prerequisite for successful implementation in clinical settings. Without robust connectivity, the usability and accuracy of such tools are compromised, limiting their effectiveness in supporting communication for CALD patients. On-device translation models offer a potential mitigation strategy because translation can be performed locally without an internet connection, reducing reliance on connectivity and potentially reducing privacy and data security risks by limiting transmission of sensitive content to external services. However, current on-device approaches may be less accurate than cloud-based translation in some scenarios, reflecting the trade‑offs of running smaller models under limited compute and memory.

It is also important to recognise that connectivity limitations also represent a potential clinical safety risk, particularly in time-sensitive situations. Interruptions to network access can result in delays, incomplete translations, or breakdowns in communication flow, potentially compromising the timely exchange of clinically relevant information. In healthcare settings, such disruptions may affect clinical decision-making, delay recognition of emerging concerns, or hinder the delivery of clear instructions to patients. In situations requiring rapid or precise communication, even brief lapses in connectivity may lead to missed or misunderstood information, increasing the risk of clinical error. Reliable network infrastructure should therefore be considered a prerequisite for safe implementation, rather than solely a usability consideration. Future studies should explore whether on-device translation models could reduce some connectivity-related disruptions observed in this study.

One of the most salient findings of this study was the discrepancy between participant-reported satisfaction and interpreter-rated translation accuracy. While a majority of participants reported that the translated communication was clear and easy to understand, interpreter assessments indicated substantial deficiencies in clarity, semantic accuracy, and fidelity to the original message. This discrepancy likely reflects fundamentally different evaluative frameworks between patients and professional interpreters. Participants may have assessed the software based on perceived understanding, conversational flow, and ability to achieve basic communication goals, particularly in a context where language barriers would otherwise limit interaction. In contrast, interpreters evaluated translations according to professional standards. Professional interpreters are trained to maintain fidelity to the source message by avoiding distortions, unjustified omissions, or additions. They also adapt their techniques to suit the interpreting mode and context. Furthermore, human interpreters employ specialised skills such as cutting in, requesting clarification or repetition, and managing turn-taking to ensure smooth communication. They are also responsible for accurately conveying the speakers’ tone and style. These nuanced practices, integral to high-quality interpreting expectations, were not feasible within the translation software’s current capabilities.

This divergence highlights the important distinction between perceived understanding and actual understanding. Patients may feel confident that communication has been successful, even when clinically relevant inaccuracies may present. This creates a potential risk of false reassurance, whereby patients believe they have correctly understood health information, but subtle mistranslations may lead to misunderstanding of clinical advice, reduced adherence, or compromised informed decision-making. This finding highlights the importance of evaluating AI translation tools using both patient-reported usability outcomes and expert linguistic or clinical safety assessments, as patient satisfaction alone may not adequately detect translation risks.

These findings also raise important medico-legal considerations. The use of AI-mediated translation introduces questions regarding accountability when communication errors occur. If a mistranslation contributes to an adverse outcome, responsibility may be challenging to determine, whether it lies with the clinician, the health service, or the technology provider.

Taken together, these considerations reinforce that, in its current form, the software should not be used as a substitute for professional interpreters in clinical settings. Instead, its use should be limited to low-risk scenarios with appropriate safeguards, including clinician awareness of potential inaccuracies, availability of interpreter oversight, and confirmation of patient understanding through techniques such as teach-back.

The primary objective of this study was to evaluate the feasibility of providing timely and seamless real-time translation during antenatal consultations for patients from CALD community. While transcription functionality was included in the software, it was not the main focus of the intervention; rather, the emphasis was on enabling effective two-way verbal communication between clinicians and patients. Importantly, professional interpreters were present throughout all consultations to monitor the process and confirm that no critical errors or risks occurred that would necessitate stopping the consultation. This oversight ensured patient safety and maintained clinical integrity. Post-consultation transcript review was conducted for evaluation purposes but was not considered essential for the immediate clinical interaction, as the primary goal was real-time communication rather than documentation.

Implications and actions needed

While professional interpreters remain the gold-standard for overcoming language barriers, findings from this study suggested that transcription and translation software might serve as a supplementary tool in carefully defined contexts, such as:

  • Communicating with patients with LEP;
  • Low-risk health care consultations, such as routine antenatal consultations, administrative interactions, or general health discussions where no complex decision-making is required;
  • Conversations using short, simple, and clear phrases;
  • Simple, structured communication, involving short phrases and clear, non-technical language;
  • Situations where interpreter access is delayed or unavailable, provided that clinician oversight is maintained;
  • Supervised environments, where clinicians remain vigilant to potential inaccuracies and confirm understanding using strategies such as teach-back.

Clear criteria are required to guide the immediate discontinuation of AI-mediated translation and escalation to professional interpreting services. These include:

  • Increased clinical complexity, such as new symptoms, complications, or changes in patient condition;
  • Repeated misunderstanding or communication breakdown, including inconsistent responses or inability to verify key information;
  • Communication of complex or sensitive content, including informed consent, discussion of risks, diagnoses, or treatment decisions;
  • Evidence of inaccurate or unclear translation, including omission, distortion, or clinician uncertainty regarding meaning;
  • High-risk situations, such as emergencies or safeguarding concerns;
  • Patient distress or reduced confidence in the communication process.

Safe implementation requires the development of clinical governance frameworks, including:

  • Clear standard operating procedures (SOPs) defining appropriate use and limitations;
  • Clinician training on the capabilities, risks, and safe use of translation technologies;
  • Defined escalation pathways to interpreter services;
  • Integration of verification strategies, such as teach-back, to confirm patient understanding to reduce the risk of miscommunication.

It is important to note that AI-powered transcription and translation software are designed to improve over time through continuous learning. As the software is exposed to a larger volume of conversational data and undergoes iterative training on diverse linguistic contexts, its performance is expected to become more accurate and contextually appropriate. Therefore, it is recommended that ongoing improvement of AI translation systems must be accompanied by continuous validation in real-world clinical settings. This should include periodic reviews of usability and performance to ensure that advances in model training and software development are reflected in practice, along with regular evaluation against clinical safety standards, monitoring for emerging errors or unintended consequences, and structured feedback from clinicians, interpreters, and patients. Establishing robust governance frameworks such as periodic audits, transparent reporting of errors, and clear accountability mechanisms will be essential to ensure that technical advancements translate into safe, reliable, and equitable clinical use.

Further, interactive AI-translation software tools should be considered in these settings. These tools could adopt human interpreter skills such as cutting in, requesting clarification or repetition, real-time error detection and managing turn-taking to improve translation quality and patient experience. Integrating these capabilities has the potential to enhance translation quality while improving the overall patient–clinician communication experience.

Future directions

Further research is needed to evaluate the software across diverse languages, dialects, and clinical settings to improve generalizability, and incorporate feedback from health care providers alongside patient and interpreter assessments. Larger-scale studies should explore integration into routine workflows and develop SOP for safe and effective use of translation technologies in maternity care and broader health care contexts.

Methodologically, future studies should incorporate blinded assessment approaches to minimize potential evaluator bias, alongside the involvement of multiple independent interpreters to improve reliability and inter-rater consistency. Studies should also include structured measures of patient comprehension, such as teach-back or scenario-based recall, to better assess the clinical safety and effectiveness of AI-mediated communication.

From a technological perspective, further development should prioritize:

  • Improved translation accuracy and contextual understanding;
  • Enhanced handling of dialectal variation and mixed-language speech;
  • Reduced latency and improved real-time responsiveness;
  • Greater system robustness, including reduced dependence on network connectivity.

Finally, ongoing validation in real-world clinical environments will be essential to ensure that improvements in model performance translate into safe, reliable, and equitable communication outcomes, particularly for CALD populations.


Conclusions

This study provides preliminary evidence on the feasibility of integrating transcription and translation software into clinical environment to support communication with CALD patients. The software demonstrated feasibility in facilitating two-way communication during low-risk antenatal consultations and was perceived as usable by a proportion of participants. These findings suggest that real-time AI-mediated translation may have potential as a supportive communication tool in settings where interpreter access is limited.

However, limitations were identified, particularly in translation accuracy, system reliability, and dependence on stable internet connectivity.

Taken together, these findings indicate that, despite demonstrating operational feasibility, the software in its current form is not suitable for independent clinical use and cannot be considered a substitute for professional interpreters. At present, software should be used only as supplementary aid in carefully controlled low-risk scenarios with appropriate clinical oversight, particularly where interpreter services are unavailable or delayed.

Future development should prioritise improvements in translation accuracy, contextual understanding, and system robustness. Further research involving larger and more diverse populations, blinded evaluation methods, and objective measures of patient comprehension is required to determine whether AI-mediated translation can meet the standards necessary for safe and effective clinical communication.


Acknowledgments

The authors would like to acknowledge the patients who participated in this study.

The authors acknowledge the use of Microsoft Copilot (GPT-based AI tool) to assist with language editing, and refinement of the manuscript. The authors take full responsibility for the content, interpretation, and conclusions presented.


Footnote

Reporting Checklist: The authors have completed the STROBE reporting checklist. Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0054/rc

Data Sharing Statement: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0054/dss

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0054/prf

Funding: This study was supported by the NSW Government through the Office of the Chief Scientist & Engineer’s Small Business Innovation & Research (SBIR) Program Grant (No. A6090083).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0054/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by South Western Sydney Local Health District Human Ethics Research Committee (2024/ETH01838). Written informed consent was obtained from all participants prior to any study procedure.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Australian Bureau of Statistics. Cultural diversity of Australia. Canberra: ABS; 2022 September 20. [cited 2025 November 13]. Available online: https://www.abs.gov.au/articles/cultural-diversity-australia
  2. Multicultural health. NSW health; December 1, 2022. [Accessed November 12, 2025]. Available online: https://www.health.nsw.gov.au/multicultural/Pages/default.aspx.
  3. Javanparast S, Naqvi SKA, Mwanri L. Health service access and utilisation amongst culturally and linguistically diverse populations in regional South Australia: a qualitative study. Rural Remote Health 2020;20:5694. [Crossref] [PubMed]
  4. Khatri RB, Assefa Y. Access to health services among culturally and linguistically diverse populations in the Australian universal health care system: issues and challenges. BMC Public Health 2022;22:880. [Crossref] [PubMed]
  5. Khatri RB, Assefa Y. Drivers of the Australian Health System towards Health Care for All: A Scoping Review and Qualitative Synthesis. Biomed Res Int 2023;2023:6648138. [Crossref] [PubMed]
  6. NSW Health. NSW Plan for Healthy Culturally and Linguistically Diverse Communities: 2019-2023; 2019. [Accessed November 12, 2025]. Available online: https://www1.health.nsw.gov.au/pds/ActivePDSDocuments/PD2019_018.pdf
  7. NSW Government. Multicultural Health Communication Service. [Accessed November 12, 2025]. Available online: https://www.mhcs.health.nsw.gov.au/about-us/cald-community
  8. South Western Sydney Local Health District: Health Snapshot. Prepared by SWSLHD Planning Unit. Updated December 2023. [Accessed November 12, 2025]. Available online: https://www.swslhd.health.nsw.gov.au/planning/content/pdf/CommunityHealthProfile/SWSLHD_Health_Snapshot_Updated_December_2023-v1.2.pdf
  9. SWSLHD Multicultural Services Implementation Plan 2021-2024. Accessed November 12, 2025. Available online: https://www.swslhd.health.nsw.gov.au/planning/pdf/MulticulturalServicesImplementationPlan.pdf
  10. Panayiotou A, Gardner A, Williams S, et al. Language Translation Apps in Health Care Settings: Expert Opinion. JMIR Mhealth Uhealth 2019;7:e11316. [Crossref] [PubMed]
  11. Harrison R, Walton M, Chitkara U, et al. Beyond translation: Engaging with culturally and linguistically diverse consumers. Health Expect 2020;23:159-68. [Crossref] [PubMed]
  12. Hwang K, Williams S, Zucchi E, et al. Testing the use of translation apps to overcome everyday healthcare communication in Australian aged-care hospital wards-An exploratory study. Nurs Open 2022;9:578-85. [Crossref] [PubMed]
  13. Hunter D, Oates R, Anderson N, et al. Validation testing of a language translation device for suitability in assisting Australian radiation therapists to communicate with Mandarin-speaking patients. Tech Innov Patient Support Radiat Oncol 2023;26:100207. [Crossref] [PubMed]
  14. Taylor B, McLean G. Exploring the use of mobile translation applications for culturally and linguistically diverse patients during medical imaging examinations in Australia - a systematic review. J Med Radiat Sci 2024;71:432-44. [Crossref] [PubMed]
  15. Hudelson P, Chappuis F. Using Voice-to-Voice Machine Translation to Overcome Language Barriers in Clinical Communication: An Exploratory Study. J Gen Intern Med 2024;39:1095-102. [Crossref] [PubMed]
  16. Silvera-Tawil D, Pocock C, Bradford D, et al. CALD Assist-Nursing: Improving communication in the absence of interpreters. J Clin Nurs 2018;27:4168-78. [Crossref] [PubMed]
  17. Freyne J, Bradford D, Pocock C, et al. Developing Digital Facilitation of Assessments in the Absence of an Interpreter: Participatory Design and Feasibility Evaluation With Allied Health Groups. JMIR Form Res 2018;2:e1. [Crossref] [PubMed]
  18. Panayiotou A, Hwang K, Williams S, et al. The perceptions of translation apps for everyday health care in healthcare workers and older people: A multi-method study. J Clin Nurs 2020;29:3516-26. [Crossref] [PubMed]
  19. Kapoor R, Corrales G, Flores MP, et al. Use of Neural Machine Translation Software for Patients With Limited English Proficiency to Assess Postoperative Pain and Nausea. JAMA Netw Open 2022;5:e221485. [Crossref] [PubMed]
  20. Abreu R, Adriatico T. Spanish for the Audiologist: Is There an App for That? Perspectives on Communication Disorders and Sciences in Culturally and Linguistically Diverse Populations 2015;22:122-8.
  21. Lee W, Khoong EC, Zeng B, et al. Evaluation of Commercially Available Machine Interpretation Applications for Simple Clinical Communication. J Gen Intern Med 2023;38:2333-9. [Crossref] [PubMed]
  22. Haith-Cooper M. Mobile translators for non-English-speaking women accessing maternity services. British Journal of Midwifery 2014;22:795-803.
  23. Turner AM, Choi YK, Dew K, et al. Evaluating the Usefulness of Translation Technologies for Emergency Response Communication: A Scenario-Based Study. JMIR Public Health Surveill 2019;5:e11171. [Crossref] [PubMed]
  24. Miller JM, Harvey EM, Bedrick S, et al. Simple patient care instructions translate best: Safety guidelines for physician use of google translate. Journal of Clinical Outcomes Management 2018;25:
  25. Bitar D, Oscarsson M. Arabic-speaking women’s experiences of communication at antenatal care in Sweden using a tablet application-Part of development and feasibility study. Midwifery 2020;84:102660. [Crossref] [PubMed]
  26. Bitar D, Oscarsson M, Hadziabdic E. Midwives’ perceptions of communication at antenatal care using a bilingual digital dialog support tool- a qualitative study. BMC Pregnancy Childbirth 2025;25:282. [Crossref] [PubMed]
  27. Johnsen H, Christensen U, Juhl M, et al. Implementing the MAMAACT intervention in Danish antenatal care: a qualitative study of non-Western immigrant women’s and midwives’ attitudes and experiences. Midwifery 2021;95:102935. [Crossref] [PubMed]
  28. Georgeou N, Schismenos S, Wali N, et al. A Scoping Review of Aging Experiences Among Culturally and Linguistically Diverse People in Australia: Toward Better Aging Policy and Cultural Well-Being for Migrant and Refugee Adults. Gerontologist 2023;63:182-99. [Crossref] [PubMed]
  29. Australian College of Midwives. National midwifery guidelines for consultation and referral. 5th ed. Sydney: Australian College of Midwives; 2026.
  30. Multicultural Health Communication Service Translation Checking Guidelines. Available online: https://www.mhcs.health.nsw.gov.au/translation-1/mhcs_translation_check_template.pdf/@@download/file
  31. Silvera-Tawil D, Pocock C, Bradford D, et al. Enabling Nurse-Patient Communication With a Mobile App: Controlled Pretest-Posttest Study With Nurses and Non-English-Speaking Patients. JMIR Nurs 2021;4:e19709. [Crossref] [PubMed]
  32. Kreienbrinck A, Hanft-Robert S, Mösko M. Usability of technological tools to overcome language barriers in health care: a scoping review protocol. BMJ Open 2024;14:e079814. [Crossref] [PubMed]
doi: 10.21037/jmai-2026-0054
Cite this article as: Chow JSF, Maurya N, Shen S, Mani S, May S, Gillard B, Hyett J, Bowkett C, Bennett L, Heys A. The feasibility of using artificial intelligence-powered translation software in lieu of human interpreters for low-risk, Arabic-speaking pregnant people in an Australian antenatal clinic. J Med Artif Intell 2026;9:62.

Download Citation