Speech-derived digital biomarkers for Alzheimer’s disease and mild cognitive impairment: a narrative review of advances, barriers, and a clinical translation roadmap
Review Article

Speech-derived digital biomarkers for Alzheimer’s disease and mild cognitive impairment: a narrative review of advances, barriers, and a clinical translation roadmap

Ya-Ting Wang1, Bao-Liang Zhong1,2,3

1Research Center for Psychological and Health Sciences, China University of Geosciences (Wuhan), Wuhan, China; 2Department of Psychiatry, Wuhan Mental Health Center, Wuhan, China; 3Hubei Clinical Research Center for Whole-Course Management of Late-Life Mental Disorders, Wuhan, China

Contributions: (I) Conception and design: Both authors; (II) Administrative support: Both authors; (III) Provision of study materials or patients: Both authors; (IV) Collection and assembly of data: Both authors; (V) Data analysis and interpretation: Both authors; (VI) Manuscript writing: Both authors; (VII) Final approval of manuscript: Both authors.

Correspondence to: Bao-Liang Zhong, MD, PhD. Department of Psychiatry, Wuhan Mental Health Center, No. 89 Gongnongbing Rd., Jiang’an District, Wuhan 430012, China; Research Center for Psychological and Health Sciences, China University of Geosciences (Wuhan), Wuhan, China; Hubei Clinical Research Center for Whole-Course Management of Late-Life Mental Disorders, Wuhan, China. Email: haizhilan@gmail.com.

Background and Objective: Alzheimer’s disease (AD) and mild cognitive impairment (MCI) require scalable, low-burden tools for early triage and longitudinal monitoring. While cerebrospinal fluid (CSF) and positron emission tomography (PET) biomarkers can confirm pathology, cost, invasiveness, and limited access constrain population-level use. Because speech production draws on semantic memory, executive control, and motor planning, AD/MCI-related neurodegeneration can yield measurable acoustic and linguistic changes. This narrative review synthesizes evidence for speech-derived digital vocal biomarkers in AD/MCI and summarizes key translational barriers and a roadmap for adoption.

Methods: We performed a narrative synthesis of English and Chinese studies (1980–January 20, 2026) from PubMed, CNKI, and benchmark repositories (e.g., DementiaBank) using terms related to AD/MCI, speech/vocal biomarkers, acoustic/linguistic features, and machine learning and natural language processing (NLP), including large language models (LLMs).

Key Content and Findings: Across elicited and naturalistic tasks (e.g., verbal fluency, picture description, recall/narration, open interviews, spoken cognitive-test responses, and telephone conversations), AD/MCI speech commonly shows slowed timing and altered pausing, reduced lexical-semantic organization, and impaired discourse coherence. Methods have progressed from engineered features with classical machine learning to deep learning and emerging self-supervised/LLM-based representations, with increasing reporting of interpretability/robustness analyses and privacy-preserving designs. Reported discrimination for AD vs. healthy controls (HCs) typically reached an area under the receiver operating characteristic curve (AUC) 0.88–0.95 (accuracy 0.82–0.96), whereas MCI vs. HCs ranged from AUC 0.87–0.99 (accuracy 0.78–0.83); direct AD vs. MCI separation was less frequently evaluated and more modest (balanced sensitivity/specificity 0.80 or accuracy 0.76). Most studies relied on internal validation, and generalizability across languages, dialects, devices, and care settings remains under-tested. Barriers include demographic and language/dialect bias, task/label heterogeneity, device/environment domain shift, and error propagation from automatic speech recognition (ASR). The roadmap prioritizes protocol harmonization; prospective multi-site external validation with subgroup reporting; clinically relevant endpoints (biomarker-anchored or adjudicated longitudinal outcomes); clinically aligned calibration/thresholds; and governance for privacy, bias, and interpretability.

Conclusions: Speech-derived biomarkers are promising, low-burden tools for triage and monitoring that complement—rather than replace—cognitive testing and biological biomarkers. Routine-care deployment will require executing the roadmap to ensure robust, equitable, and clinically meaningful performance.

Keywords: Alzheimer’s disease (AD); mild cognitive impairment (MCI); speech biomarkers; digital biomarkers; artificial intelligence (AI)


Received: 08 February 2026; Accepted: 09 March 2026; Published online: 16 March 2026.

doi: 10.21037/jmai-2026-1-0032


Introduction

Dementia—particularly Alzheimer’s disease (AD), accounting for approximately 50–75% of cases—and its prodromal stage, mild cognitive impairment (MCI), remains a major public health challenge (1-4). By 2050, the global number of individuals living with dementia is projected to be 130.8–175.9 million, with substantial impacts on healthcare systems and caregivers (5). Current diagnostic workflows typically combine clinical assessment, neuropsychological testing [e.g., Mini-Mental State Examination (MMSE), Montreal Cognitive Assessment (MoCA)], and biological or imaging biomarkers. Cerebrospinal fluid (CSF) measures of β-amyloid (Aβ) and phosphorylated tau (p-tau), alongside positron emission tomography (PET) imaging of Aβ/tau deposition, support pathological confirmation but are constrained by cost, invasiveness, and limited access in many settings (6,7). These constraints contribute to delayed recognition and limit opportunities for early intervention and trial referral, motivating scalable, low-burden prescreening tools such as speech-derived digital biomarkers (8,9).

Speech is a ubiquitous, low-burden behavior that reflects integrated function across semantic memory, executive control, attention, and motor planning (10). Subtle disruptions in these systems—often preceding overt functional decline—can manifest as quantifiable changes in both acoustic properties (e.g., speech rate, pitch stability, pause dynamics) and linguistic/discourse organization (e.g., lexical diversity, semantic coherence, idea density) (11). Despite rapid progress in artificial intelligence (AI)-enabled speech analysis, translation is limited by heterogeneity in tasks and labels, domain shift across devices and environments, and insufficient prospective validation against biomarker-confirmed endpoints.

Several recent reviews have summarized the rapidly growing literature on speech/voice-derived digital biomarkers for cognitive impairment, covering datasets, methods, and overall diagnostic promise, while noting heterogeneity and limited real-world validation (12-14). However, practical guidance for deployment remains incomplete—particularly regarding standardization for replication, robustness across languages/dialects and recording conditions, handling automatic speech recognition (ASR)–related errors, clinically anchored validation endpoints/designs, and governance for privacy, bias/fairness, and interpretability amid foundation-model and large language model (LLM) workflows.

Building on this landscape, this narrative review synthesizes the historical evolution and recent advances in speech-derived biomarkers for AD/MCI, spanning foundational definitions, mechanistic rationale, data collection and analysis pipelines, and evidence organized by clinical use (screening, diagnosis/staging, and prognosis). Importantly, this review advances the field by providing a translation-oriented synthesis that links end-to-end workflows (from acquisition and preprocessing through modeling and evaluation, including self-supervised speech representations and LLMs) with a consolidated view of deployment barriers and a practical roadmap for responsible implementation in clinical care, primary-care settings, and community screening programs. We present this article in accordance with the Narrative Review reporting checklist (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0032/rc).


Methods

We conducted a narrative synthesis of studies on speech/voice-derived digital biomarkers for AD/MCI, searching PubMed and CNKI and supplementing with benchmark/dataset portals and challenge websites (e.g., DementiaBank/ADReSS/ADReSSo) and reference lists from database inception to January 20, 2026 (search date: January 20, 2026; full strategy in Table 1). We included English- or Chinese-language empirical studies and eligible reviews/meta-analyses evaluating speech/voice-derived features or models for AD/MCI screening/diagnosis/prediction, and excluded non-AD/MCI populations without relevant subgroup analyses, non-speech/voice studies, purely theoretical work without empirical evaluation, and reports lacking sufficient methodological detail. Two authors independently screened titles/abstracts and then full texts against predefined criteria, resolving disagreements by discussion and consensus; we then purposively selected representative original studies and recent systematic reviews/meta-analyses (January 1, 2020 to January 20, 2026) for translation-focused synthesis.

Table 1

Summary of the literature search strategy

Items Specification
Date of search January 20, 2026
Databases and other sources searched PubMed; CNKI; benchmark/dataset portals and challenge websites (e.g., DementiaBank/ADReSS/ADReSSo); reference lists of eligible reviews and key papers
Search terms used Keywords in title/abstract/keywords included combinations of: “Alzheimer’s disease”, “mild cognitive impairment”, “dementia”, “speech”, “voice”, “vocal biomarker”, “speech biomarker”, “digital biomarker”, “acoustic”, “prosody”, “pause”, “jitter”, “shimmer”, “natural language processing”, “large language model”, “machine learning”, “deep learning”, “verbal fluency”, “picture description”, and “story recall”
Timeframe From database inception to January 20, 2026
Inclusion and exclusion criteria Inclusion criteria: English or Chinese studies (original research, reviews/meta-analyses, benchmark/challenge reports) evaluating speech/voice-derived features or models for AD/MCI screening/diagnosis/prediction, including remote/community/primary-care deployment studies when available
Exclusion criteria: Non-AD/MCI populations without relevant subgroup analysis; studies unrelated to speech/voice signals; purely theoretical papers without empirical evaluation; studies lacking sufficient methodological detail for interpretation.
Selection process Two authors independently screened titles/abstracts and then full texts against predefined criteria; disagreements were resolved by discussion and consensus

AD, Alzheimer’s disease; MCI, mild cognitive impairment.


Definition of speech-derived digital biomarkers

Speech-derived digital biomarkers are objective, quantifiable features derived from audio speech signals and/or transcripts that correlate with cognitive status or pathology-related changes (15). Unlike subjective clinician descriptors (e.g., “disfluent speech”), speech-based digital biomarkers are extracted computationally through signal processing and natural language processing (NLP) and can capture early, subtle deviations in speech planning, semantic retrieval, and motor control (16). Some studies suggest certain acoustic instabilities and linguistic impoverishment may precede overt memory complaints by several years in at-risk populations (15,17).

Compared with CSF and PET biomarkers, speech-based assessment is non-invasive, comparatively low cost, and scalable via consumer devices—supporting repeated, longitudinal monitoring and potential population-level triage (14,18) (Table 2). These strengths are especially relevant in primary care and underserved regions where advanced imaging is limited and where frequent reassessment is clinically desirable but operationally difficult with traditional biomarkers (8,19). Nevertheless, it is crucial to acknowledge the limitations highlighted in Table 2: unlike the single-point biological specificity of CSF or PET, speech metrics are highly sensitive to inter-individual variability (e.g., educational level, dialect) and environmental factors. Therefore, their optimal utility may lie in tracking intra-individual changes over time rather than serving as a standalone cross-sectional diagnostic.

Table 2

Comparative analysis of speech-derived digital biomarkers vs. traditional Alzheimer’s disease biomarkers (CSF/PET)

Feature    Speech-based vocal biomarkers    Traditional biomarkers (CSF/PET)
Primary target    Phenotypic/functional: detects downstream cognitive-linguistic & motor deficits    Pathological/biological: directly measures proteinopathy (Aβ, Tau)
Invasiveness & cost    Low: non-invasive, scalable, remote (smartphone-based)    High: invasive (lumbar puncture) or requires radioactive tracers; expensive
Accessibility    High: can be deployed in community/home settings    Low: restricted to specialized clinical centers
Time-point sensitivity    Limitations (longitudinal): due to high inter-individual variability (e.g., education, dialect), often requires longitudinal monitoring (intra-individual change) for optimal accuracy    Strength (cross-sectional): highly effective at a single time-point for establishing biological diagnosis
Confounding factors    High sensitivity to context: affected by education level, mood (depression/anxiety), environmental noise, and recording hardware    Low environmental impact: results are generally stable regardless of patient’s education or mood state
Specific utility    Screening & monitoring: ideal for “Digital Triage” and detecting subtle daily fluctuations    Confirmatory diagnosis: the “Gold Standard” for biological confirmation of disease etiology

Aβ, β-amyloid; CSF, cerebrospinal fluid; PET, positron emission tomography.

To ensure conceptual clarity, it is essential to distinguish speech-derived digital biomarkers from conventional clinical assessments and general speech technologies. Unlike traditional neuropsychological tests [e.g., verbal fluency tasks (VFTs)] that rely on manual, macroscopic scoring of accuracy—which can be subject to inter-rater variability—digital biomarkers utilize automated algorithms to capture high-dimensional, sub-perceptual cues, such as millisecond-level pause dynamics or vocal micro-tremors. Furthermore, distinct from standard ASR systems that aim to filter out noise to maximize transcription accuracy, speech biomarker analysis specifically targets these “imperfections” (e.g., hesitations, repetitions, and articulatory slurring) as the primary signals of interest reflecting underlying neurodegenerative pathology.

Speech-based vocal biomarkers are commonly grouped into three interacting classes: acoustic, linguistic, and paralinguistic (20). Acoustic features are generally less language-dependent than lexical measures, but remain sensitive to task design, recording conditions, and language-specific prosody (21). Representative markers include reduced speech rate, increased within-phrase pauses and response latency, elevated jitter/shimmer, reduced harmonics-to-noise ratio, and altered spectral/cepstral patterns such as Mel Frequency Cepstral Coefficients (MFCCs), which can proxy articulatory precision (22).

Linguistic biomarkers are derived from transcripts and reflect semantic memory and executive organization (23). Common findings include reduced lexical diversity [e.g., type-token ratio (TTR)], increased pronoun/generic-word use (“it,” “thing”), simplified syntax (shorter utterances, reduced syntactic depth), and diminished discourse coherence and idea density, increasingly quantified via contextual embeddings [e.g., bidirectional encoder representations from transformers (BERT)-like models] (24).

Paralinguistic biomarkers capture non-verbal communicative signals (affective prosody and pragmatics), including reduced pitch variability/monotony, disrupted conversational turn-taking, and increased repetition—potentially reflecting frontal–subcortical dysfunction and social-cognitive impairment (25).


Historical evolution of vocal biomarker research for AD/MCI

From signal processing to machine learning

Over the past several decades, speech-based biomarker research for AD/MCI has evolved from descriptive clinical studies to increasingly automated AI pipelines, in step with advances in speech/language processing and machine learning (18,26). In the early exploration phase (1980s–2010), studies primarily relied on manual transcription and clinician-annotated behavioral assessments to characterize perceptible speech and language changes, such as anomia, circumlocution, and tangentiality, often in very small cohorts (18,27). Basic acoustic measurements (e.g., speech rate and pitch) became possible with early tools such as Praat, but analyses were labor-intensive and rarely standardized across sites or tasks (26,28). Subsequent work has begun to relate speech-timing abnormalities to structural brain changes (including temporal-lobe–related regions), while semi-standardized elicitation paradigms—most notably the Cookie Theft picture description from the BDAE—have been widely used to elicit connected speech in dementia research (29,30). Overall, many studies have faced translation barriers—small samples, heterogeneous protocols, and limited standardization/comparability—constraining robust model-based validation and clinical implementation (18,26).

The feature engineering and traditional machine-learning phase (2010–2019) shifted the field toward quantification and early standardization, aiming to convert clinical impressions into reproducible acoustic and linguistic markers that could support prediction (26,31). Toolchains matured for systematic feature extraction—OpenSMILE enabling standardized acoustic sets [e.g., extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS)] and Praat supporting voice-quality measures such as jitter and shimmer—while NLP pipelines [e.g., tokenization, part-of-speech (POS) tagging] enabled transcript-derived features including type–token ratio and syntactic proxies (26,32). With these engineered features, classical classifiers [especially support vector machines (SVMs), alongside random forests and logistic regression] became widely adopted due to their strong performance on small datasets (26,31). The period also saw early systematic syntheses of acoustic and linguistic markers in AD, alongside increasing use of shared corpora (e.g., DementiaBank) that facilitated more comparable benchmarking across models (26,33). Yet despite improved measurability, this era remained limited by dependence on handcrafted features, frequent reliance on human transcription, and insufficient cross-cohort external validation—leaving generalizability as a recurring concern (26,31).

Recent advances: the era of deep learning and LLMs

The deep learning phase (2020–2023) accelerated automation by reducing reliance on manual feature design and enabling end-to-end representation learning from raw audio or spectrograms (12,34). Architectures, such as convolutional neural networks (CNNs; for spectral patterns) and long short-term memory (LSTM; for temporal dynamics), as well as fusion models combining acoustic and linguistic channels, often improved performance on benchmark datasets and highlighted the value of multimodal integration (12,35). Standardized resources expanded (e.g., ADReSSo and SpeechDx-related efforts), supporting more consistent evaluation and facilitating model comparisons across research groups (36,37). This phase also began to push speech biomarkers beyond “classification demos” toward clinical-relevance claims, including validation of composite speech outcomes as potential trial endpoints and demonstrations that adding complementary modalities (e.g., gait) can substantially improve discrimination in some settings (38,39). At the same time, the field confronted new bottlenecks: deep models were often less interpretable, data remained relatively scarce and English-skewed, and high benchmark accuracy did not necessarily translate into robust performance under real-world recording conditions (12,34).

Most recently, the LLM and clinical translation phase (2024–2026) has emphasized real-world workflow integration and the governance frameworks needed for deployment (40). LLMs and embedding approaches (e.g., transformer representations) support deeper semantic/discourse characterization (e.g., coherence) and have been explored for marker discovery and data augmentation, though clinical validity depends on careful control of synthetic-data artifacts (41,42). In parallel, privacy-preserving strategies such as federated learning address the sensitivity of voice data, while remote and ambient collection modalities (smartphone-based testing, passive monitoring) make longitudinal assessment increasingly practical (43,44). Reported milestones in this era include large harmonized speech datasets linked to biomarker measures, including speech datasets linked to biomarker measures (including PET/CSF) and conversational agents capable of administering remote cognitive assessments, collectively signaling a shift from “model performance” to “system deployment” (45). The central challenge now is to convert these technical gains into prospective, multi-site, biomarker-anchored evidence with interpretable outputs, robustness across subgroups/recording conditions, and regulation-ready documentation, so that speech biomarkers can complement frontline screening pathways rather than remain confined to benchmarks (46,47).

Figure 1 depicts the evolution of speech analysis over time, illustrating the shift from manual feature engineering to automated representation learning.

Figure 1 Timeline of technological evolution in speech-based biomarkers for AD and MCI. The evolution of speech analysis is depicted along a chronological axis, illustrating the shift from manual feature engineering to automated representation learning. (I) Early phase (1980s–2010): focus on descriptive statistics of acoustic features (e.g., pitch, pause duration) using small clinical datasets. (II) Machine learning phase (2010–2019): adoption of standardized feature sets (e.g., eGeMAPS) and classical classifiers (SVM, random forest). (III) Deep learning & LLM era (2020–present): the current frontier, characterized by end-to-end deep neural networks and pre-trained LLMs that capture high-dimensional linguistic and acoustic representations. AD, Alzheimer’s disease; BERT, bidirectional encoder representations from transformers; eGeMAPS, extended Geneva Minimalistic Acoustic Parameter Set; LLMs, large language models; MCI, mild cognitive impairment; SVM, support vector machine.

Mechanisms linking speech changes to AD/MCI

Speech changes observed in AD and MCI are unlikely to be random artifacts; rather, they can be interpreted as functional readouts of region-specific neurodegeneration, network-level disconnection, and domain-general cognitive decline that jointly shape language production and speech-motor control. Because spoken language requires coordinated semantic access, executive planning, attentional allocation, and fine-grained motor timing, even subtle disruption to these systems can yield measurable deviations in acoustic timing, voice quality stability, and lexical-discourse organization—often before deficits become obvious in routine clinical conversation. Mechanistically, current evidence converges on three complementary layers: (I) AD pathophysiology disrupting speech-related neural networks; (II) associations between speech features and canonical AD pathological biomarkers; and (III) cognitive-psychological processes that explain why certain tasks reliably elicit discriminative speech signatures.

Neural mechanisms: AD pathophysiology and speech-language networks

AD-related amyloid-β and tau accumulation and spread have been linked to progressive disruption of large-scale structural and functional brain networks, including association and limbic systems that support higher-order cognition (48). Consistent with the typical anatomical progression into temporal association cortex, AD can compromise lexical-semantic representations access, yielding early word-finding difficulty (anomia) and related semantic retrieval errors (49). In connected-speech tasks, these lexical-semantic vulnerabilities are often reflected in lower lexical diversity (e.g., reduced TTR) and increased reliance on less specific referring expressions (e.g., producing “the thing you sit on” instead of “chair”), with severity-related gradients (50). Functional imaging further supports a temporo-semantic account: poorer naming and semantic fluency performance has been associated with hypometabolism in inferior temporal (and broader temporal) regions in AD (51).

As pathology extends into frontal association systems, deficits in executive functions—including working memory, inhibition, and planning—become increasingly relevant to speech output (29). These impairments can manifest as reduced phonemic fluency, simplified syntactic structures (shorter, less complex sentences), increased fillers, and degraded discourse coherence or topic maintenance (52). Comparisons with primary progressive aphasia (PPA) variants further underscore the network logic: syndromes dominated by focal temporal atrophy (e.g., semantic variant PPA) show prominent naming/semantic deficits, whereas AD more often exhibits mixed temporal-frontal involvement and therefore combined lexical-syntactic/discourse vulnerabilities (53). In other words, the pattern of speech impairment can be viewed as a behavioral footprint of which language and control networks are most affected (54).

A third, increasingly discussed mechanism concerns the motor-cognitive interface of speech (55). Motor speech deficits were historically considered late-stage features of AD, yet recent work suggests that subtle disruption of speech motor control and feedback integration may appear earlier than previously assumed (56). In this framework, reduced integrity of the speech motor control network and the somatosensory/auditory feedback monitoring systems can yield measurable micro-instabilities—such as elevated jitter and shimmer—and increased variability in speech rate as the system compensates for a “noisier” internal model of articulatory control (56). This provides a plausible mechanistic bridge between cortical neurodegeneration and acoustic markers that may be detectable even when linguistic content appears superficially intact (57).

Brain science evidence: linking speech markers to pathological biomarkers

A critical step toward clinical credibility is demonstrating that speech-derived features track disease-relevant biology, not only nonspecific aging or psychosocial factors (58). Several studies have reported that combinations of acoustic and linguistic markers can predict amyloid PET positivity in cognitively unimpaired or mildly impaired individuals with moderate discrimination (e.g., AUCs in the 0.71–0.83 range in both cohorts with and without cognitive impairment), suggesting that subtle alterations in prosody/voice quality and timing may align with preclinical amyloid burden (18). Task choice appears to matter: narrative paradigms (e.g., personal storytelling) have been reported to capture early changes such as reduced pitch range or altered prosodic dynamics, consistent with the hypothesis that early network inefficiencies may be expressed as constrained vocal variability (58).

Parallel evidence links language and timing measures to tau-related processes (59). Linguistic markers (e.g., pause dynamics, lexical diversity) have been associated with CSF p-tau levels in AD-spectrum phenotypes, including language-led presentations such as logopenic PPA, supporting a model in which tau accumulation in fronto-temporal circuitry disrupts neural timing and retrieval efficiency, thereby increasing pausing and reducing informational content (60). Importantly, these associations remain heterogeneous across cohorts and pipelines; thus, they should be interpreted as convergent support for biological relevance rather than definitive proof of specificity (61). Nevertheless, the emerging alignment between speech markers and amyloid/tau measures strengthens the rationale for using speech as a low-cost prescreener to prioritize confirmatory—but expensive—biomarker testing (62).

Psychological mechanisms: speech as an integrative probe of cognition

From a cognitive-psychological perspective, speech production is a sensitive “stress test” of multiple cognitive domains, which explains why AD/MCI often shows characteristic patterns across tasks (63). Semantic memory supports word retrieval and concept specificity; its erosion contributes to “empty speech”, overuse of generic terms/pronouns, and reduced lexical diversity (27). Episodic memory deficits, particularly salient in spontaneous narrative tasks, can yield disjointed discourse, reduced event specificity, and weakened global coherence (64). Meanwhile, declines in executive control (working memory for syntactic planning; inhibitory control for maintaining topic and suppressing irrelevant associations) provide a mechanistic account for tangentiality, shortened sentences, and increased fillers—phenomena that may become more pronounced as conversational demands rise (52).

This cognitive account also clarifies why measurement context can influence observed speech features (33). For example, intensive prompting or scaffolding by an examiner may partially compensate for executive-control deficits and artificially inflate apparent fluency in structured settings—an “observer effect” that highlights the importance of standardized administration and careful reporting of prompting behavior (52). Finally, cognitive load offers a principled way to amplify disease-relevant vulnerabilities: dual-task paradigms (e.g., speaking while walking) tax shared attentional and working-memory resources and can exacerbate slowing and pause frequency, which helps explain why some studies report improved discrimination under dual-task conditions compared with single-task speech alone (e.g., higher AUC in certain cohorts) (52). Collectively, these observations support a task-informed view of vocal biomarkers: speech features are not merely static traits, but context-sensitive outputs that reveal how neurodegenerative changes interact with cognitive demand.

In summary, mechanistic evidence supports a coherent chain from AD pathology and network disruption (temporal-semantic, frontal-executive, and motor-feedback systems) to task-dependent acoustic and linguistic signatures, with growing—though still incomplete—biological anchoring to amyloid and tau measures. This framework motivates careful task selection, standardized administration, and biomarker-confirmed validation as key steps toward clinically reliable speech-based screening tools.

Figure 2 presents a conceptual framework showing how AD and MCI pathology in specific neural/cognitive networks translates hierarchically into measurable speech-derived digital biomarkers across acoustic, paralinguistic, and linguistic feature dimensions.

Figure 2 Conceptual framework of speech-derived digital biomarkers in AD and MCI. The diagram illustrates a hierarchical mapping from brain pathology to observable speech manifestations and speech-derived digital features. AD pathology affects distinct neural-cognitive networks (motor cortex for motor planning; frontal–subcortical/limbic circuits for executive and social-cognitive control; temporal lobes for semantic memory), leading to semantic, executive, and motor speech deficits. These deficits can be quantified by three clusters of speech-based digital biomarkers: acoustic (e.g., jitter, pause duration), paralinguistic (e.g., flat affect/monotony, disrupted turn-taking), and linguistic features (e.g., reduced vocabulary, simplified syntax, reduced semantic coherence). AD, Alzheimer’s disease; MCI, mild cognitive impairment.

Acquisition, processing, and modeling pipeline of speech-based vocal biomarkers

Speech can be collected in clinics (high control), at home (high ecological validity), or in community settings (scalable screening) (65). Task choice critically shapes which cognitive systems are engaged and which biomarkers emerge (66). Picture description tasks, such as “Cookie Theft”, standardize stimulus content and elicit semi-structured spontaneous speech that probes naming, syntax, and narrative organization (33). VFTs probe semantic memory (semantic fluency) and executive control (phonemic fluency) (67). Narrative recall or story-based tasks may sensitively detect subtle discourse and episodic-memory vulnerabilities and have been incorporated into remote protocols [e.g., Storyteller within Alzheimer’s Disease Neuroimaging Initiative 4 (ADNI)-related work] (65). Dual-task paradigms (e.g., speaking while walking) can amplify deficits by increasing cognitive load and may improve discrimination in some settings, especially when combined with gait features (68).

Preprocessing typically includes voice activity detection, denoising, and amplitude/sampling normalization to reduce device/environment variability (69). For transcript-based biomarkers, ASR enables scalable text extraction but introduces a key dependency: ASR error rates often increase in older or cognitively impaired speech, and errors can distort lexical/syntactic features if unreported or unmodeled (70). Best practice is to report ASR performance (e.g., word error rate), evaluate sensitivity to transcription errors, and consider hybrid pipelines that combine robust acoustics with transcript-based semantics (71). To avoid optimistic bias, evaluation should ensure participant-level splits (no speaker overlap) and, ideally, external testing across sites/devices (72).

Earlier approaches relied on handcrafted acoustic (e.g., eGeMAPS) and linguistic features (e.g., TTR, POS ratios, syntactic depth), paired with traditional classifiers (SVM, random forests, logistic regression) (73). Deep learning enabled automatic feature learning from spectrograms or raw waveforms and improved performance on benchmark datasets and multimodal fusion settings (74). More recently, LLM-style embeddings enable richer modeling of semantics, discourse structure, and topic drift, and have been explored for interpretable marker discovery and data augmentation—though clinical validity depends on careful evaluation for distribution shift, label noise, and synthetic artifacts when synthetic data are used (41,75).


Evidence base for screening, diagnosis, prediction, and multimodal fusion

Speech-derived biomarkers for AD/MCI have expanded rapidly, but reported performance is highly heterogeneous across elicitation tasks (e.g., picture description, verbal fluency, recall, routine interviews), reference standards (clinical diagnosis vs. cognitive screening scores vs. biomarker-confirmed endpoints), and validation designs (internal cross-validation vs. held-out testing vs. external multi-site testing). Accordingly, we summarize evidence by intended clinical use: screening (MCI vs. HCs; broader cognitive impairment vs. HCs), diagnosis (AD vs. HCs), staging (AD vs. MCI), prediction, and multimodal fusion, emphasizing validation rigor and deployability rather than headline accuracy alone.

Screening (MCI vs. HCs; cognitive impairment vs. HCs)

Across peer-reviewed studies, screening-oriented contrasts (MCI or early impairment vs. HCs) often achieve moderate-to-high discrimination when models integrate timing/pause dynamics with lexical/semantic organization, but performance varies substantially with cohort and endpoint definition. For example, app-based multi-task batteries in multi-site cohorts can reach screening-relevant discrimination (e.g., HCs vs. MCI/probable AD AUC =0.83) while maintaining interpretable feature sets, supporting feasibility for low-burden prescreening workflows (76).

Community-scale brief conversational protocols also report strong AUCs when using modern self-supervised acoustic representations (e.g., Wav2Vec2-derived voice vectors), suggesting that short, natural speech can support scalable triage (77).

In biomarker-adjacent or carefully characterized research cohorts, story recall-based speech and transcript modeling can yield very strong held-out discrimination for early cognitive impairment (e.g., AUC =0.945 using engineered speech/language features and up to 0.988 using an end-to-end transcript model), though these results should be interpreted in light of cohort composition and generalizability (78).

Importantly, many studies remain limited by small sample sizes and predominantly internal validation, underscoring the need for external testing and subgroup robustness reporting before real-world screening claims are made.

Diagnosis (AD vs. HCs)

For AD vs. HCs, performance is often higher than MCI screening—consistent with larger effect sizes at later disease stages and more salient disruptions in fluency, rate, and discourse content. Task design matters: higher memory-load speech tasks (e.g., recall) can outperform lower-demand tasks in telephone-based protocols, supporting the “cognitive stress test” framing for speech biomarkers (79).

Deep learning pipelines using short speech responses embedded within routine cognitive workflows (e.g., MMSE-derived spoken answers transformed into program inputs) also report strong AD-HCs discrimination under cross-validation (80).

In structured multi-task settings, automated feature extraction with classical ML can also reach high AD-HCs accuracy in recent studies (e.g., HCs vs. probable AD with an AUC of 0.90 in a speech-language multi-task protocol), illustrating that interpretable timing and prosodic markers retain value even without end-to-end models (76).

Staging (AD vs. MCI)

Distinguishing AD from MCI is often more challenging than AD vs HCs because MCI is a symptomatic predementia stage and a heterogeneous syndrome with variable trajectories (including reversion), which increases diagnostic uncertainty and can introduce label noise (81). Accordingly, AD-MCI performance is frequently lower and more variable than AD-HCs in applied settings; for example, in a memory-clinic multi-task speech protocol (e.g., counting backward, sentence repetition, picture description, verbal fluency) using a classical SVM-based pipeline, König et al. (2015) reported classification accuracy of 87% for AD vs. HCs and 80% for AD vs. MCI (82).

Studies that explicitly evaluate MCI vs. mild AD (mAD) using spontaneous speech and ASR-supported acoustic features combined with linguistic features show that fusing acoustic and linguistic information can improve discrimination versus acoustics alone (e.g., MCI vs. mAD accuracy increasing from 76.0% with extended acoustic features alone to 80.0% with the fusion of acoustic and full linguistic features, as well as 78.0% with the fusion of acoustic and semantic linguistic features in specific configurations) (83).

Because staging is sensitive to diagnostic boundaries and reference standards, future work should prioritize biomarker-anchored endpoints or well-defined longitudinal outcomes to reduce label ambiguity and improve clinical interpretability.

Prediction (longitudinal change; time-to-onset/conversion)

Longitudinal prediction remains less mature than cross-sectional classification. Systematic syntheses consistently highlight that the field still lacks broad cohort-style repeated testing and standardized endpoints needed to prove reliability for within-person tracking and conversion prediction in real-world settings (13,14,57,84).

Multimodal fusion (speech + other behavioral/biomarkers)

Multimodal fusion is a recurring strategy to increase robustness and specificity by triangulating partially independent disease signals. In a three-way classification framework (HCs/MCI/AD), combining speech with gait and drawing can substantially improve staging performance compared with speech alone (e.g., multi-class AUC rising from 0.91 to 0.98 in one cohort), consistent with complementary information across behavioral modalities (39).

At the review level, an explainable-AI-focused synthesis reports a graded trend where multimodal models often show higher AUC ranges than unimodal acoustic systems, but also concludes that formal pooling is frequently infeasible due to heterogeneity in tasks, datasets, and evaluation protocols (84).

Table 3 lists representative studies using speech- and language-derived digital biomarkers for AD and MCI detection and summarizes the performance of the corresponding AI-based models.

Table 3

Representative studies using speech-derived digital biomarkers for AD and MCI detection/classification

Study Setting Participants Cognitive vocal tasks Modality Artificial intelligence algorithm Validation Accuracy indicators
Talar et al., 2024 (76) Community 312 MCI, 272 probable AD, 417 HCs Multi-task battery (reading, phonation, image description, story recall, naming) Acoustic + language features XGBoost Internal: holdout HCs vs. (MCI/pAD): AUC =0.83; HCs vs. pAD: AUC =0.90
Rezaii et al., 2025 (78) Memory clinics 120 MCI, 68 HCs Delayed recall speech + transcripts Engineered speech + language and end-to-end transcript model XGBoost and RoBERTa-base Internal: holdout MCI vs. HCs: AUC =0.945 (engineered) and AUC =0.988 (end-to-end)
König et al., 2015 (82) Memory clinics 26 early AD, 23 MCI, 15 HCs Four brief vocal tasks including picture description, semantic fluency Acoustic + prosodic features Support vector machine Internal: holdout Equal sensitivity and specificity: 0.79 (MCI vs. HCs), 0.87 (AD vs. HCs), 0.80 (MCI vs. AD)
Gosztolya et al., 2019 (83) Memory clinics 25 mild AD, 25 MCI, 25 HCs Film recall + narrative tasks Acoustic + linguistic (ASR-supported) features + age +sex + education Support vector machine Internal: holdout Accuracy: 0.78 (MCI vs. HCs), 0.82 (AD vs. HCs), 0.76 (AD vs. MCI)
Kiyoshige et al., 2025 (77) Day care centers and communities 526 MCI, 935 HCs 3-min open interview Acoustic + prosodic features + age + sex + education Depp neural network Internal: holdout MCI vs. HCs: AUC =0.89
Ding et al., 2024 (85) Communities 100 MCI, 100 HCs Speech from neuropsychological tests Nonsemantic acoustic features Random forest Internal: resampling MCI vs. HCs: AUC =0.87
Bae et al., 2023 (79) A medical center and communities 45 mild-to-moderate AD and 44 HCs 3 speech tasks with varying memory loads Acoustic features A weighted voting classifier combining random forest and logistic regression Internal: holdout AD vs. HCs: AUC =0.88 (high memory load task)
Ahn et al., 2023 (80) A university-affiliated hospital and communities 40 AD, 40 HCs Spoken MMSE answers Acoustic features Convolutional neural network: Resnet50 Internal: holdout AD vs. HCs: AUC 0.95
Yamada et al., 2021 (39) A psychiatric outpatient at a university-affiliated hospital and communities 26 AD, 45 MCI, 47 HCs 5 tablet-based speech cognitive tasks + gait + drawing Acoustic + linguistic + prosodic features + gait + drawing features A support vector machine with a radial basis function kernel Internal: resampling Accuracy (speech-only): 0.93 (AD vs. HCs), 0.83 (MCI vs. HCs); Accuracy (speech + gait + drawing): 1.00 (AD vs. HCs), 0.90 (MCI vs. HCs)
Yang et al., 2022 (86) Recordings of calls to a health insurer help line 651 AD, 1,018 HCs Unstructured natural telephone conversations Acoustic + prosodic + linguistic features Multi-layered perceptron Internal: resampling AD vs. HCs: accuracy 0.96

AD, Alzheimer’s disease; ASR, automatic speech recognition; AUC, area under the curve; HCs, healthy controls; MCI, mild cognitive impairment; MMSE, Mini-Mental State Examination; pAD, probable Alzheimer’s disease.


Evidence synthesis from recent systematic reviews and real-world illustration

Evidence synthesis from recent systematic reviews

Beyond individual studies, recent systematic reviews converge on two robust conclusions: (I) integrating acoustic and linguistic features tends to outperform either modality alone, and (II) reported metrics remain difficult to meta-analyze cleanly because datasets, labels, tasks, and validation protocols vary widely. For example, a 2025 systematic review of NLP-based approaches reported higher average performance for combined pipelines (average accuracy 87%; AUC =0.89) than for linguistic-only (83%; AUC =0.85) or acoustic-only approaches (80%; AUC =0.82), while emphasizing substantial heterogeneity across included studies (24). In parallel, an acoustic-focused systematic review/meta-analysis of continuous-speech markers quantified group differences in rate- and interruption-related measures for AD vs. HCs, supporting the biological plausibility of acoustic change, but did not yield universal pooled classification metrics for MCI or AD-MCI staging (13).

Taken together, these syntheses suggest that multimodal integration is promising, while also highlighting the need for standardized datasets, task designs, and validation/reporting to improve comparability and generalizability.

Real-world illustration: Wuhan community/primary-care pilot (unpublished)

To illustrate how the above synthesis may translate into practice-oriented settings, our group conducted a preliminary feasibility study in Wuhan community and primary-care environments (unpublished data; not peer-reviewed) using a brief semantic VFT. We enrolled 67 older adults with early-stage AD, 55 with MCI, and 84 cognitively HCs, and recorded speech during a 1-minute animal-naming task administered in ≤3 minutes with minimal instructions and no specialized equipment. The task was well tolerated (98.5% completion; no withdrawals due to task difficulty). We processed speech using a standardized pipeline and extracted acoustic/paralinguistic features spanning timing/fluency and voice-quality cues (e.g., speech rate, pause measures, perturbation metrics), alongside linguistic features (e.g., lexical diversity and frequency-based measures); response latency was treated as an interactional timing measure. Using subject-level evaluation with 10-fold cross-validation and conventional classifiers (support vector machines; random forests), early fusion of acoustic and linguistic features yielded sensitivities of 0.911 (early AD vs. HCs) and 0.884 (MCI vs. HCs) at a prespecified operating point, with corresponding specificities of 0.923 and 0.846, respectively. In an exploratory post hoc analysis, several MCI participants with nominally “normal” VFT word counts (≥15 words/min) nonetheless showed atypical speech patterns captured by the model, suggesting potential incremental sensitivity beyond paper-based scoring.

These findings are intended as a real-world illustration of feasibility and potential signal utility, and should be interpreted cautiously pending preregistered, prospective multi-site validation with standardized protocols, subgroup-robustness reporting, and external replication.


Discussion

Limitations of this narrative review

As a narrative (rather than fully systematic) review, this work does not include an exhaustive search or a formal methodological quality/risk-of-bias assessment (e.g., Prediction model Risk Of Bias ASsessment Tool) for every included study. Reported performance ranges—including the AUC values summarized in Table 3—therefore reflect substantial heterogeneity in tasks, cohorts, endpoints, and evaluation protocols, and should be interpreted cautiously as descriptive benchmarks rather than definitive estimates of real-world clinical utility. The underlying literature also varies in design rigor and reporting transparency, with potential biases arising from small single-site samples, heterogeneous diagnostic labeling (clinical vs biomarker-confirmed), limited independent external validation, and evaluation choices that may inflate performance (e.g., non-independent cross-validation, leakage, or selective metric reporting). Finally, some evidence on commercial implementations appears in non-peer-reviewed formats (e.g., conference proceedings or white papers) and should be treated as preliminary, prioritizing independent, preregistered, multi-site prospective validation with standardized protocols and robust subgroup reporting.

Key barriers to translation

Several field-wide barriers must be addressed to enable clinical adoption.

Generalizability and bias

A dominant barrier is limited external validity due to demographic, linguistic, and cultural imbalance in available datasets, with many corpora overrepresenting English-speaking, highly educated populations (87,88). This is not only an equity issue but a scientific validity issue: language structure (e.g., tonal vs. non-tonal), dialect/accent, education, hearing status, and sociocultural narrative norms can shift feature distributions and confound disease effects (88,89). Without subgroup reporting and targeted adaptation, models risk encoding population- or device-specific artifacts rather than disease-relevant signals (46).

Task and label heterogeneity

Cross-study comparability is undermined by heterogeneity in elicitation tasks (picture description vs. free conversation vs. recall) and diagnostic labeling rigor (clinical diagnosis vs. biomarker-confirmed AD pathology) (12,31). Because different tasks recruit different cognitive processes, a biomarker that is discriminative in one paradigm may not replicate in another, contributing to heterogeneity in meta-analytic estimates (18). Progress requires harmonized task batteries and transparent reporting of prompts, timing, and scoring procedures.

Domain shift in real-world audio

Models trained on controlled clinical recordings frequently degrade in noisy home or community environments due to background noise, microphone differences, compression, and conversational structure (36). If robustness is not explicitly modeled and evaluated, systems may learn “site/device signatures” and fail at deployment. Practical translation therefore, requires evaluation on real-world recordings, domain adaptation strategies, and reporting of performance under varied acoustic conditions.

Interpretability, calibration, and clinical utility

“Black-box” models hinder clinician trust and regulatory evaluation. Explainability should be clinically meaningful (e.g., stable markers tied to plausible cognitive mechanisms) rather than post hoc visualization alone (90). Equally important, studies must move beyond accuracy to demonstrate calibration, decision thresholds aligned with referral pathways, and explicit modeling of harm/benefit trade-offs for false positives/negatives in the intended setting (primary care vs. specialist clinic) (91).

Privacy, governance, and regulation

Voice is inherently identifying, and naive de-identification can distort acoustic features central to many biomarkers (92). Privacy-preserving learning (e.g., federated learning) is promising but requires rigorous governance, auditability, and security considerations (93). From a regulatory standpoint, systems intended for medical decision-making may be evaluated as software as a medical device (SaMD), making prospective multi-site clinical validation and clear intended-use statements essential.

To bridge the gap between research and practice, recent initiatives offer specific lessons. Regarding harmonized protocols, the DementiaBank and TalkBank repositories demonstrate the value of standardized task instructions (e.g., the Cookie Theft picture description); however, future protocols must also standardize recording hardware to mitigate the channel effects seen when comparing studio-quality recordings with telephone-based data in cohorts like the Framingham Heart Study. For multi-site external validation, recent benchmark competitions (e.g., the ADReSS challenges) reveal that models achieving >90% accuracy on matched datasets often suffer significant performance drops (e.g., <60%) when applied to independent, unmatched cohorts, underscoring the critical need for domain-invariant training. Finally, defining clinical decision thresholds requires moving beyond binary classification; similar to plasma biomarkers, speech scores should be calibrated against biological endpoints (e.g., amyloid/tau status) to define specific cut-offs that maximize sensitivity for early screening while maintaining high specificity for diagnostic referral.

Roadmap and future directions

To move speech-based vocal biomarkers from promising prototypes to clinically dependable tools, the field needs a roadmap that progresses from method standardization, to robust evidence generation, to biological anchoring, and finally to deployment-ready governance—while keeping the intended clinical setting (often primary care and community screening) as the organizing principle (Table 4).

Table 4

Strategic roadmap for speech-derived biomarker translation: from barriers to clinical adoption

Current barriers (challenges)    Strategic roadmap (solutions) Target clinical utility
Bias & generalizability (demographic imbalance, dialects)    Inclusive data & adaptation (federated learning, cross-lingual training) Equitable screening (population-level filter)
Task heterogeneity (varied protocols/hardware)    Harmonization (standardized hardware & tasks, e.g., DementiaBank) Scalability (remote monitoring)
Label reliability (clinical labels only)    Biological anchoring (validation vs. CSF/PET biomarkers) Precision diagnosis (gatekeeper for biomarkers)
“Black Box” nature (lack of explainability)    Interpretable AI (clinically meaningful features, XAI) Clinician trust (adoption in primary care)
Privacy & regulation (data sensitivity)    Governance & SaMD (privacy-preserving architecture, clear intended use) Regulatory approval (medical device certification)

AI, artificial intelligence; CSF, cerebrospinal fluid; PET, positron emission tomography.

First, translation cannot begin without harmonized protocols, because reproducibility is the foundation of both meta-analysis and regulatory-grade evidence. In practice, this means standardizing prompts, task duration, microphone guidance, and preprocessing steps, and publishing protocol details with enough specificity to support true replication across sites and languages (36). Without this layer, performance differences across studies remain uninterpretable: it becomes unclear whether models are capturing disease-related signals or simply reflecting task design choices and recording artifacts.

Once protocols are aligned, the next requirement is prospective, multi-site validation, which is the most direct test of external validity. Rather than optimizing for in-sample accuracy, studies should prespecify endpoints, include subgroup analyses, and report calibration and decision-curve metrics alongside AUC/accuracy to demonstrate clinical usefulness under realistic referral pathways. This shift—from “can the model classify?” to “does the model support better decisions?”—is essential for deployment in settings where false positives and false negatives carry different downstream consequences.

However, external validity alone is insufficient if the target is AD-specific screening. It is important to acknowledge that current literature relies heavily on clinical diagnoses, with a paucity of direct validation against fluid or imaging biomarkers. Consequently, the field should increasingly prioritize biomarker-anchored endpoints—amyloid/tau status (e.g., PET or CSF) or rigorously adjudicated longitudinal conversion outcomes—to strengthen pathological specificity and reduce the risk that models capture nonspecific effects (e.g., normal aging or depression-related speech changes) rather than Alzheimer’s pathology (18). Anchoring speech markers to biology is also the most convincing way to align with contemporary AD diagnostic frameworks and to support clinical adoption as a triage layer for expensive confirmatory tests.

In parallel, models should be engineered with robustness-by-design because real-world audio is the norm, not the exception. Training and evaluation need to reflect realistic noise, device variability, and conversational structure, with explicit quantification of sensitivity to ASR errors and careful consideration of hybrid acoustic–linguistic pipelines when transcripts are unreliable. Robustness should be treated as a primary outcome rather than a post hoc add-on, since deployment typically fails at precisely this interface between controlled data and uncontrolled settings.

As systems approach real-world use, deployable governance becomes a technical and ethical prerequisite, not an administrative afterthought. Privacy-preserving approaches such as federated learning can reduce the need to centralize identifiable voice data, but they must be paired with clear consent models, explicit intended-use statements, defined failure modes, and monitoring plans for drift and bias over time. These elements are increasingly intertwined with regulatory expectations for SaMD and with clinician trust.

Finally, the roadmap should remain grounded in where speech biomarkers are most likely to create immediate value: community and primary-care screening. Brief, feasible tasks—particularly semantic VFTs (e.g., 1-minute animal naming)—offer an operational advantage (low burden, minimal training, smartphone-compatible) that aligns naturally with large-scale outreach. Yet feasibility alone is not enough; future studies should explicitly quantify the incremental value of speech-derived features beyond conventional VFT word counts and standard screeners (e.g., MoCA), using prespecified thresholds, clinically meaningful outcomes, and subgroup robustness testing in real-world recording conditions. Demonstrating incremental benefit is the clearest way to justify adoption in busy frontline settings where any additional tool must earn its place in the workflow.

This sequence—standardize, validate externally, anchor biologically, build robustness, govern for deployment, and prove incremental value in community settings—provides a coherent pathway for converting speech-based biomarkers into scalable, credible clinical instruments.

To maximize clinical utility, we propose integrating speech biomarkers via a tiered “digital triage” pathway where they complement, rather than replace, existing diagnostic workflows. In the first tier, speech tools deployed on consumer smartphones can serve as a scalable, high-sensitivity filter for population-level screening, identifying at-risk individuals remotely who might otherwise be missed. Subsequently, within primary care, these biomarkers act as an adjunct to standard cognitive screeners (e.g., MoCA or MMSE), where features like reduced semantic density can detect subtle deficits often masked by the ceiling effects of traditional paper-and-pencil tests. Ultimately, speech analysis functions as a cost-effective gatekeeper for biological confirmation, enriching the pre-test probability of AD pathology to help clinicians prioritize patients for invasive or expensive diagnostics (e.g., Amyloid-PET or CSF assays), thereby optimizing healthcare resource allocation.


Conclusions

Speech-derived digital vocal biomarkers have progressed from descriptive observations to AI-enabled digital measures that can capture both acoustic-motor control changes and higher-order linguistic and discourse disruptions relevant to AD/MCI. The literature supports their potential as scalable, low-burden tools for triage screening and longitudinal monitoring, particularly when protocols elicit spontaneous language under meaningful cognitive demand and when evaluation is grounded in external validation and clinically aligned decision thresholds. However, widespread adoption depends on solving generalizability and domain-shift challenges, reducing task/label heterogeneity through harmonization, strengthening validation against biomarker-confirmed endpoints, and implementing interpretable, privacy-preserving, and regulation-ready workflows. Preliminary, unpublished data from a Wuhan community/primary-care pilot further support the feasibility and acceptability of a brief semantic VFT for scalable screening in older adults and suggest that speech-derived measures may capture impairment signals not reflected by word-count scoring alone. With rigorous prospective evidence and responsible governance, speech biomarkers are well positioned to complement conventional cognitive testing and biological biomarkers within multimodal precision screening and monitoring pathways for cognitive decline.


Acknowledgments

During the preparation of this manuscript, the authors used Doubao 1.5 Pro solely for language editing (e.g., improving readability and correcting grammar). No AI-generated content was used to generate or interpret data, and all scientific conclusions remain the authors’own. The authors reviewed and edited the output and take full responsibility for the content.


Footnote

Reporting Checklist: The authors have completed the Narrative Review reporting checklist. Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0032/rc

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0032/prf

Funding: This study was supported by the Wuhan Municipal Health Commission and Bureau of Science and Technology Innovation of Wuhan Municipality (No. WX23A99), National Natural Science Foundation of China (No. 71774060), and the Young Top Talent Program in Public Health from Health Commission of Hubei Province (No. EWEITONG[2021]74, PI: B-LZ).

Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0032/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Alzheimer’s Disease International. Dementia facts & figures 2024. Available online: https://www.alzint.org/about/dementia-facts-figures/. Accessed on March 7, 2026.
  2. Wang ZQ, Fei L, Xu YM, et al. Prevalence and correlates of suspected dementia in older adults receiving primary healthcare in Wuhan, China: A multicenter cross-sectional survey. Front Public Health 2022;10:1032118. [Crossref] [PubMed]
  3. Zhi N, Ren R, Qi J, et al. The China Alzheimer Report 2025. Gen Psychiatr 2025;38:e102020. [Crossref] [PubMed]
  4. Huang L, Li Q, Lu Y, et al. Consensus on rapid screening for prodromal Alzheimer’s disease in China. Gen Psychiatr 2024;37:e101310. [Crossref] [PubMed]
  5. Estimation of the global prevalence of dementia in 2019 and forecasted prevalence in 2050: an analysis for the Global Burden of Disease Study 2019. Lancet Public Health 2022;7:e105-25. [Crossref] [PubMed]
  6. Leuzy A, Bollack A, Pellegrino D, et al. Considerations in the clinical use of amyloid PET and CSF biomarkers for Alzheimer’s disease. Alzheimers Dement 2025;21:e14528. [Crossref] [PubMed]
  7. Jiang H, Escamilla S, Beestrum M, et al. Cost-effectiveness of testing biofluid biomarkers to diagnose Alzheimer’s disease: a systematic review. Eur J Health Econ 2025; Epub ahead of print. [Crossref]
  8. Angioni D, Delrieu J, Hansson O, et al. Blood Biomarkers from Research Use to Clinical Practice: What Must Be Done? A Report from the EU/US CTAD Task Force. J Prev Alzheimers Dis 2022;9:569-79. [Crossref] [PubMed]
  9. Wang J, Zeng HY, Gao JX, et al. Enhancing Dementia Awareness and Screening, and Reducing Stigmatizing Attitudes towards Dementia in Urban China: The Role of Opinion Leader Intervention in Community-Dwelling Older Adults. Alpha Psychiatry 2025;26:38857. [Crossref] [PubMed]
  10. Tremblay P, Deschamps I, Dick A. Neuromotor Organization of Speech Production. The Oxford Handbook of Speech Production. Oxford: Oxford University Press; 2019.
  11. Mueller KD, Hermann B, Mecollari J, et al. Connected speech and language in mild cognitive impairment and Alzheimer’s disease: A review of picture description tasks. J Clin Exp Neuropsychol 2018;40:917-39. [Crossref] [PubMed]
  12. Ding K, Chetty M, Noori Hoshyar A, et al. Speech based detection of Alzheimer’s disease: a survey of AI techniques, datasets and challenges. Artif Intell Rev 2024;57:325.
  13. Saeedi S, Hetjens S, Grimm MOW, et al. Acoustic Speech Analysis in Alzheimer’s Disease: A Systematic Review and Meta-Analysis. J Prev Alzheimers Dis 2024;11:1789-97. [Crossref] [PubMed]
  14. Jafari Z, Andrew MK, Rockwood KJ. Diagnostic utility of speech-based biomarkers in mild cognitive impairment: a systematic review and meta-analysis. Age Ageing 2025;54:afaf316. [Crossref] [PubMed]
  15. Hajjar I, Okafor M, Choi JD, et al. Development of digital voice biomarkers and associations with cognition, cerebrospinal biomarkers, and neural representation in early Alzheimer’s disease. Alzheimers Dement (Amst) 2023;15:e12393. [Crossref] [PubMed]
  16. Tavabi N, Stück D, Signorini A, et al. Cognitive Digital Biomarkers from Automated Transcription of Spoken Language. J Prev Alzheimers Dis 2022;9:791-800. [Crossref] [PubMed]
  17. Eyigoz E, Mathur S, Santamaria M, et al. Linguistic markers predict onset of Alzheimer’s disease. EClinicalMedicine 2020;28:100583. [Crossref] [PubMed]
  18. Fristed E, Skirrow C, Meszaros M, et al. Leveraging speech and artificial intelligence to screen for early Alzheimer’s disease and amyloid beta positivity. Brain Commun 2022;4:fcac231. [Crossref] [PubMed]
  19. Cay G, Pfeifer VA, Lee M, et al. Harnessing Speech-Derived Digital Biomarkers to Detect and Quantify Cognitive Decline Severity in Older Adults. Gerontology 2024;70:429-38. [Crossref] [PubMed]
  20. Sharafeldeen A, Keowen J, Shaffie A. Machine Learning Approaches for Speech-Based Alzheimer’s Detection: A Comprehensive Survey. Computers 2025;14:36.
  21. Campbell EL, Mesía RY, Docío-Fernández L, et al. Paralinguistic and linguistic fluency features for Alzheimer’s disease detection. Comput Speech Lang 2021;68:101198.
  22. Hason L, Krishnan S. Spontaneous speech feature analysis for alzheimer’s disease screening using a random forest classifier. Front Digit Health 2022;4:901419. [Crossref] [PubMed]
  23. Chapin K, Clarke N, Garrard P, et al. A finer-grained linguistic profile of Alzheimer’s disease and Mild Cognitive Impairment. J Neurolinguistics 2022;63:101069. PMID. [Crossref]
  24. Shankar R, Bundele A, Mukhopadhyay A. A Systematic Review of Natural Language Processing Techniques for Early Detection of Cognitive Impairment. Mayo Clin Proc Digit Health 2025;3:100205. [Crossref] [PubMed]
  25. Nasreen S, Hough J, Purver M. Detecting Alzheimer’s Disease Using Interactional and Acoustic Features from Spontaneous Speech. Proc Interspeech 2021;2021:1962-6. [Crossref]
  26. de la Fuente Garcia S, Ritchie CW, Luz S. Artificial Intelligence, Speech, and Language Processing Approaches to Monitoring Alzheimer’s Disease: A Systematic Review. J Alzheimers Dis 2020;78:1547-74. [Crossref] [PubMed]
  27. Taler V, Phillips NA. Language performance in Alzheimer’s disease and mild cognitive impairment: a comparative review. J Clin Exp Neuropsychol 2008;30:501-56. [Crossref] [PubMed]
  28. Boersma P, Weenink D. Praat: doing phonetics by computer. Version 6.2.06, retrieved 23 January 2026. Available online: https://www.praat.org, 1992-2022. Accessed on March 7, 2026.
  29. De Looze C, Dehsarvi A, Crosby L, et al. Cognitive and Structural Correlates of Conversational Speech Timing in Mild Cognitive Impairment and Mild-to-Moderate Alzheimer’s Disease: Relevance for Early Detection Approaches. Front Aging Neurosci 2021;13:637404. [Crossref] [PubMed]
  30. Cummings L. Describing the cookie theft picture Sources of breakdown in Alzheimer’s dementia. Pragmatics Soc 2019;10:153-176.
  31. Petti U, Baker S, Korhonen A. A systematic literature review of automatic Alzheimer’s disease detection from speech and language. J Am Med Inform Assoc 2020;27:1784-97. [Crossref] [PubMed]
  32. Eyben F, Scherer KR, Schuller BW, et al. The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for voice research and affective computing. IEEE Trans Affect Comput 2016;7:190-202.
  33. Lanzi AM, Saylor AK, Fromm D, et al. DementiaBank: Theoretical Rationale, Protocol, and Illustrative Analyses. Am J Speech Lang Pathol 2023;32:426-38. [Crossref] [PubMed]
  34. Shi M, Cheung G, Shahamiri SR. Speech and language processing with deep learning for dementia diagnosis: A systematic review. Psychiatry Res 2023;329:115538. [Crossref] [PubMed]
  35. Mahajan P, Baths V. Acoustic and Language Based Deep Learning Approaches for Alzheimer’s Dementia Detection From Spontaneous Speech. Front Aging Neurosci 2021;13:623607. [Crossref] [PubMed]
  36. Luz S, Haider F, de la Fuente S, Fromm D, MacWhinney B. Detecting Cognitive Decline Using Speech Only: The ADReSSo Challenge. Proc Interspeech 2021;2021:3780-4. [Crossref]
  37. PR-Newswire. The ADDF’s Diagnostics Accelerator unveils SpeechDx dataset and new partnership to catalyze breakthroughs in speech-based biomarkers 2023. Available online: https://www.prnewswire.com/news-releases/the-addfs-diagnostics-accelerator-unveils-speechdx-dataset-and-new-partnership-to-catalyze-breakthroughs-in-speech-based-biomarkers-302570634.html. Accessed on March 7, 2026
  38. Spilka MJ, Xu M, Toth B, et al. Robustness and generalizability of a speech-based digital biomarker derived from recordings of the Clinical Dementia Rating (CDR) interview. Alzheimer’s Dement 2024;20:e094733.
  39. Yamada Y, Shinkawa K, Kobayashi M, et al. Combining Multimodal Behavioral Data of Gait, Speech, and Drawing for Classification of Alzheimer’s Disease and Mild Cognitive Impairment. J Alzheimers Dis 2021;84:315-27. [Crossref] [PubMed]
  40. Cohen IG, Ritzman J, Cahill RF. Ambient Listening-Legal and Ethical Issues. JAMA Netw Open 2025;8:e2460642. [Crossref] [PubMed]
  41. Zolnour A, Azadmaleki H, Haghbin Y, et al. LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data. Front Artif Intell 2025;8:1669896. [Crossref] [PubMed]
  42. Arora A, Wagner SK, Carpenter R, et al. The urgent need to accelerate synthetic data privacy frameworks for medical research. Lancet Digit Health 2025;7:e157-60. [Crossref] [PubMed]
  43. Wang T, Du Y, Gong Y, et al. Applications of Federated Learning in Mobile Health: Scoping Review. J Med Internet Res 2023;25:e43006. [Crossref] [PubMed]
  44. Gregory S, Harrison J, Herrmann J, et al. Remote data collection speech analysis in people at risk for Alzheimer’s disease dementia: usability and acceptability results. Front Dement 2023;2:1271156. [Crossref] [PubMed]
  45. Pacheco-Lorenzo MR, Anido-Rifón LE, Fernández-Iglesias MJ, et al. Will senior adults accept being cognitively assessed by a conversational agent? a user-interaction pilot study. Appl Intell 2024;54:7897-912.
  46. Kim M, Youn YC, Won Y, et al. Domain generalization for voice-based cognitive impairment detection. BMC Med Inform Decis Mak 2025;25:450. [Crossref] [PubMed]
  47. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
  48. Yu M, Sporns O, Saykin AJ. The human connectome in Alzheimer disease - relationship to biomarkers and genetics. Nat Rev Neurol 2021;17:545-63. [Crossref] [PubMed]
  49. Altmann LJ, McClung JS. Effects of semantic impairment on language use in Alzheimer’s disease. Semin Speech Lang 2008;29:18-31. [Crossref] [PubMed]
  50. Kavé G, Dassa A. Severity of Alzheimer’s disease and language features in picture descriptions. Aphasiology 2018;32:27-40.
  51. Melrose RJ, Campa OM, Harwood DG, et al. The neural correlates of naming and fluency deficits in Alzheimer’s disease: an FDG-PET study. Int J Geriatr Psychiatry 2009;24:885-93. [Crossref] [PubMed]
  52. Kempler D, Goral M. Language and Dementia: Neuropsychological Aspects. Annu Rev Appl Linguist 2008;28:73-90. [Crossref] [PubMed]
  53. Seeley WW, Crawford RK, Zhou J, et al. Neurodegenerative diseases target large-scale human brain networks. Neuron 2009;62:42-52. [Crossref] [PubMed]
  54. Drzezga A. The Network Degeneration Hypothesis: Spread of Neurodegenerative Patterns Along Neuronal Brain Networks. J Nucl Med 2018;59:1645-8. [Crossref] [PubMed]
  55. Guenther FH. Cortical interactions underlying the production of speech sounds. J Commun Disord 2006;39:350-65. [Crossref] [PubMed]
  56. Themistocleous C, Eckerström M, Kokkinakis D. Voice quality and speech fluency distinguish individuals with Mild Cognitive Impairment from Healthy Controls. PLoS One 2020;15:e0236009. [Crossref] [PubMed]
  57. Martínez-Nicolás I, Llorente TE, Martínez-Sánchez F, et al. Ten Years of Research on Automatic Voice and Speech Analysis of People With Alzheimer’s Disease and Mild Cognitive Impairment: A Systematic Review Article. Front Psychol 2021;12:620251. [Crossref] [PubMed]
  58. van den Berg RL, de Boer C, Zwan MD, et al. Digital remote assessment of speech acoustics in cognitively unimpaired adults: feasibility, reliability and associations with amyloid pathology. Alzheimers Res Ther 2024;16:176. [Crossref] [PubMed]
  59. Nevler N, Ash S, Irwin DJ, et al. Validated automatic speech biomarkers in primary progressive aphasia. Ann Clin Transl Neurol 2019;6:4-14. [Crossref] [PubMed]
  60. Cho S, Cousins KAQ, Shellikeri S, et al. Lexical and Acoustic Speech Features Relating to Alzheimer Disease Pathology. Neurology 2022;99:e313-22. [Crossref] [PubMed]
  61. Robin J, Harrison JE, Kaufman LD, et al. Evaluation of Speech-Based Digital Biomarkers: Review and Recommendations. Digit Biomark 2020;4:99-108. [Crossref] [PubMed]
  62. König A, Linz N, Baykara E, et al. Screening over Speech in Unselected Populations for Clinical Trials in AD (PROSPECT-AD): Study Design and Protocol. J Prev Alzheimers Dis 2023;10:314-21. [Crossref] [PubMed]
  63. Mueller KD. Using discourse as a measure of early cognitive decline associated with Alzheimer’s disease biomarkers. In: Kong APH, editor. Spoken discourse impairments in the neurogenic populations. Heidelberg: Springer; 2023:53-63.
  64. Kong AP, Cheung RTH, Wong GHY, et al. Spoken discourse in episodic autobiographical and verbal short-term memory in Chinese people with dementia: the roles of global coherence and informativeness. Front Psychol 2023;14:1124477. [Crossref] [PubMed]
  65. Skirrow C, Meepegama U, Weston J, et al. Storyteller in ADNI4: Application of an early Alzheimer’s disease screening tool using brief, remote, and speech-based testing. Alzheimers Dement 2024;20:7248-62. [Crossref] [PubMed]
  66. Bose A, Dutta M, Dash NS, et al. Importance of Task Selection for Connected Speech Analysis in Patients with Alzheimer’s Disease from an Ethnically Diverse Sample. J Alzheimers Dis 2022;87:1475-81. [Crossref] [PubMed]
  67. Troyer AK, Moscovitch M, Winocur G. Clustering and switching as two components of verbal fluency: evidence from younger and older healthy adults. Neuropsychology 1997;11:138-46. [Crossref] [PubMed]
  68. Montero-Odasso MM, Sarquis-Adamson Y, Speechley M, et al. Association of Dual-Task Gait With Incident Dementia in Mild Cognitive Impairment: Results From the Gait and Brain Study. JAMA Neurol 2017;74:857-65. [Crossref] [PubMed]
  69. Shah Z, Sawalha J, Tasnim M, et al. Learning Language and Acoustic Models for Identifying Alzheimer’s Dementia From Speech. Front Comput Sci 2021;3:624659.
  70. Heitz T, Meftah S, Dupont S, et al. The influence of automatic speech recognition on linguistic features for Alzheimer’s disease detection: A systematic assessment. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). Torino, Italy: ELRA-ICCL; 2024:6790-802.
  71. Rohanian M, Hough J, Purver M. Alzheimer’s Dementia Recognition Using Acoustic, Lexical, Disfluency and Speech Pause Features Robust to Noisy Inputs. Proc Interspeech 2021;2021:3820-4. [Crossref]
  72. Luz S, Haider F, de la Fuente S, et al. Alzheimer’s Dementia Recognition Through Spontaneous Speech: The ADReSS Challenge. Proc Interspeech 2020;2020:2172-6. [Crossref]
  73. Valsaraj A, Madala I, Garg N, et al. Alzheimer’s dementia detection using acoustic & linguistic features and pre-trained BERT. arXiv 2021. doi: 10.48550/arXiv.2109.11010.
  74. Ilias L, Askounis D, Psarras J. Detecting dementia from speech and transcripts using transformers. Comput Speech Lang 2023;79:101485.
  75. Agbavor F, Liang H. Predicting dementia from spontaneous speech using large language models. PLOS Digit Health 2022;1:e0000168. [Crossref] [PubMed]
  76. Talkar T, Charles S, Krantsevich C, Kawabata K. Detection of Cognitive Impairment And Alzheimer’s Disease Using a Speech- and Language-Based Protocol. Proc Interspeech 2024;2024:3025-9. [Crossref]
  77. Kiyoshige E, Ogata S, Kwon N, et al. Developing and testing AI-based voice biomarker models to detect cognitive impairment among community dwelling adults: a cross-sectional study in Japan. Lancet Reg Health West Pac 2025;59:101598. [Crossref] [PubMed]
  78. Rezaii N, Wong B, Aisen P, et al. Voiceprints of cognitive impairment: analyzing digital voice for early detection of Alzheimer’s and related dementias. NPJ Dement 2025;1:35. [Crossref] [PubMed]
  79. Bae M, Seo MG, Ko H, et al. The efficacy of memory load on speech-based detection of Alzheimer’s disease. Front Aging Neurosci 2023;15:1186786. [Crossref] [PubMed]
  80. Ahn K, Cho M, Kim SW, et al. Deep Learning of Speech Data for Early Detection of Alzheimer’s Disease in the Elderly. Bioengineering (Basel) 2023;10:1093. [Crossref] [PubMed]
  81. Langa KM, Levine DA. The diagnosis and management of mild cognitive impairment: a clinical review. JAMA 2014;312:2551-61. [Crossref] [PubMed]
  82. König A, Satt A, Sorin A, et al. Automatic speech analysis for the assessment of patients with predementia and Alzheimer’s disease. Alzheimers Dement (Amst) 2015;1:112-24. [Crossref] [PubMed]
  83. Gosztolya G, Vincze V, Tóth L, et al. Identifying mild cognitive impairment and mild Alzheimer’s disease based on spontaneous speech using ASR and linguistic features. Comput Speech Lang 2019;53:181-97.
  84. Shankar R, Goh Z, Devi F, et al. A systematic review of explainable artificial intelligence methods for speech-based cognitive decline detection. NPJ Digit Med 2025;8:724. [Crossref] [PubMed]
  85. Ding H, Lister A, Karjadi C, et al. Detection of Mild Cognitive Impairment From Non-Semantic, Acoustic Voice Features: The Framingham Heart Study. JMIR Aging 2024;7:e55126. [Crossref] [PubMed]
  86. Yang K, O’Connell H. Alzheimer’s Model Performance (Whitepaper): Canary Speech; 2022. Available online: https://canaryspeech.com/wp-content/uploads/2022/08/Canary-Speech-Alzheimers-Model-Performance.pdf. Accessed on March 7, 2026.
  87. Guo Y, Li C, Roan C, et al. Crossing the ‘Cookie Theft’ Corpus Chasm: Applying what BERT Learns from Outside Data to the ADReSS Challenge Dementia Detection Task. Front Comput Sci 2021;3:642517. [Crossref] [PubMed]
  88. García AM, de Leon J, Tee BL, et al. Speech and language markers of neurodegeneration: a call for global equity. Brain 2023;146:4870-9. [Crossref] [PubMed]
  89. Koenecke A, Nam A, Lake E, et al. Racial disparities in automated speech recognition. Proc Natl Acad Sci U S A 2020;117:7684-9. [Crossref] [PubMed]
  90. World Health Organization. Ethics and governance of artificial intelligence for health: WHO guidance. World Health Organization; 2021. Available online: https://iris.who.int/handle/10665/341996. Accessed on March 7, 2026
  91. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making 2006;26:565-74. [Crossref] [PubMed]
  92. Patino J, Tomashenko N, Todisco M, et al. Speaker Anonymisation Using the McAdams Coefficient. Proc Interspeech 2021;2021:1099-1103. [Crossref]
  93. Eden R, Chukwudi I, Bain C, et al. A scoping review of the governance of federated learning in healthcare. NPJ Digit Med 2025;8:427. [Crossref] [PubMed]
doi: 10.21037/jmai-2026-1-0032
Cite this article as: Wang YT, Zhong BL. Speech-derived digital biomarkers for Alzheimer’s disease and mild cognitive impairment: a narrative review of advances, barriers, and a clinical translation roadmap. J Med Artif Intell 2026;9:45.

Download Citation