Multi-cohort transcriptomic analysis and ensemble machine learning identify IGF1R and SPP1 as candidate diagnostic genes and nominate compound leads for Alzheimer’s disease
Highlight box
Key findings
• A leakage-controlled multi-cohort framework integrating differential expression analysis, weighted gene co-expression network analysis, protein-protein interaction ranking, and 127 prespecified machine-learning model-feature combinations identified IGF1R and SPP1 as reproducible Alzheimer’s disease (AD)-associated candidate diagnostic genes.
• The final random forest model achieved an area under the curve (AUC) of 0.957 (95% confidence interval: 0.922–0.984), sensitivity of 0.978, and specificity of 0.936; four compounds showed predicted docking energies below −9.0 kcal/mol, including one at −10.5 kcal/mol.
What is known and what is new?
• AD is molecularly heterogeneous, and current biomarkers are limited by invasiveness, cost, availability, and variable reproducibility across cohorts; transcriptomic machine-learning studies are further prone to optimistic bias arising from data leakage.
• This study adds a leakage-controlled, externally validated workflow combining network biology, ensemble machine learning, SHapley Additive exPlanations interpretation, immune-infiltration analysis, and single-cell transcriptomics. IGF1R and SPP1 maintained AUC values above 0.70 in all three external validation cohorts and were associated with distinct neuronal-metabolic and inflammatory biological contexts.
What is the implication, and what should change now?
• IGF1R and SPP1 should be prioritized for prospective multicenter validation in accessible biospecimens, including blood and cerebrospinal fluid, and for mechanistic characterization, rather than being treated as validated biomarkers. The nominated compounds require biochemical, cellular, animal-model, pharmacokinetic, and clinical validation before any therapeutic interpretation or application.
Introduction
Alzheimer’s disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia worldwide. Its pathological hallmarks include amyloid-β deposition, tau pathology, neuroinflammation, synaptic loss, and metabolic dysregulation (1). AD is increasingly recognized as a biologically heterogeneous condition involving interconnected alterations across neuronal, glial, immune, vascular, and metabolic systems. This biological heterogeneity is reflected clinically, as early symptoms may overlap with normal aging, vascular cognitive impairment, depression, or other neurodegenerative disorders, making early and accurate diagnosis challenging (2-5). Diagnostic uncertainty may delay disease recognition and reduce the opportunity for timely intervention and therapeutic decision-making.
Amyloid and tau positron emission tomography (PET), cerebrospinal fluid (CSF) biomarkers, and emerging blood-based markers are now incorporated into biomarker-centered diagnostic frameworks, greatly improving the biological characterization of AD (6,7). However, PET imaging is expensive and not widely available in many clinical settings, CSF collection is invasive, and blood-based markers require further standardization across assay platforms, populations, and care settings (8,9). Furthermore, single modality biomarker panels may not capture AD biology fully as there is coordinated molecular dysregulation across neuronal, glial, immune, vascular, and metabolic systems (10,11). Reproducible transcriptomic signatures can provide additional value to current diagnostic strategies by capturing disease-associated molecular perturbations that are not fully reflected by conventional biomarker modalities.
Public transcriptomic datasets enable investigation of disease-associated molecular alterations across independent cohorts without requiring new primary data generation (12,13). Integrating multiple datasets introduces challenges such as batch effects, platform heterogeneity, and small-sample instability, requiring careful cohort partitioning to ensure that validation datasets remain independent of feature selection and model development (14). A multi-layer analytical approach combining differentially expressed gene (DEG), co-expression network analysis, and protein-protein interaction (PPI) prioritization may improve the robustness of candidate-gene selection by requiring convergent support from independent analytical perspectives (15).
Machine-learning models provide a systematic framework for evaluating whether network-prioritized candidate genes carry reproducible diagnostic information across independent cohorts (16). Evaluating multiple algorithm families, including tree-based ensemble, gradient boosting, and regularized regression models, under a unified internal cross-validation (CV) framework can provide a more comprehensive assessment of diagnostic signals than reliance on a single classifier (17,18). Model interpretability is also essential in high-dimensional omics-based diagnostic studies (19). SHapley Additive exPlanations (SHAP) can clarify the relative contribution of individual genes to model predictions, while immune deconvolution and single-cell transcriptomic analysis can provide complementary biological context for interpreting candidate genes (20).
In this study, we applied an integrative multi-cohort transcriptomic framework to nominate candidate diagnostic genes for AD. DEGs and disease-associated co-expression modules were identified in the discovery cohort, and PPI network analysis was used to prioritize network-supported candidates. These candidates were then evaluated using 127 prespecified machine-learning model-feature combinations, with final hub genes selected by integrating SHAP-derived model contribution, bootstrap stability, PPI-rank stability, and cross-cohort reproducibility across three external validation cohorts. Biological relevance was further examined through immune infiltration analysis and single-cell transcriptomic profiling. Finally, drug-gene enrichment, ADMET (absorption, distribution, metabolism, excretion, and toxicity) profiling, and molecular docking were used to generate exploratory compound-prioritization hypotheses for future experimental investigation.
Methods
Study design
This study was conducted as a secondary analysis of publicly available AD transcriptomic datasets. As shown in Figure 1, the workflow connected transcriptomic discovery, network-based gene prioritization, diagnostic model development, biological contextualization, and exploratory compound prioritization. AD-associated transcriptional changes were identified in the discovery cohort and integrated with disease-related co-expression modules. Genes supported by differential expression and co‑expression network evidence were mapped to a PPI network for candidate-gene prioritization. The diagnostic performance of prioritized genes was evaluated using machine-learning models developed in the training cohort and assessed in external validation. Immune infiltration analysis and single-cell transcriptomic analysis were used to characterize the biological context of the final candidate hub genes. Drug-gene enrichment, drug-likeness assessment, ADMET prediction, and molecular docking generated compound-prioritization hypotheses for future experimental studies.
All molecular discovery steps—including differential expression analysis (DEA), co-expression network construction, PPI-based candidate ranking, model training, hyperparameter optimization, and SHAP-based interpretation—were performed exclusively within the training cohort. GSE110226, GSE28146, and GSE29378 served as external validation cohorts and were accessed only after the analytical framework had been fully specified within the training cohort. All datasets were publicly available and de-identified; no primary clinical specimens or experimental models were generated in this study. All identified candidates are interpreted as computationally nominated hypotheses requiring experimental and clinical validation. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was conducted as a secondary analysis of publicly available, de-identified transcriptomic datasets. No new human subjects research, animal experiments, clinical specimens, or patient-identifiable data were generated or collected; therefore, additional institutional ethics approval and informed consent were not required.
Data sources and cohort definition
Bulk transcriptomic datasets were retrieved from the Gene Expression Omnibus (GEO) database (21). GSE5281 and GSE66333 were assigned to the training cohort. GSE110226, GSE28146, and GSE29378 were assigned to external validation. GSE147047 was used for single-cell transcriptomic analysis.
The training cohort was used for DEA, weighted gene co-expression network analysis (WGCNA), PPI network construction, candidate-gene ranking, model training, hyperparameter optimization, and SHAP-based model interpretation. The external validation cohorts were used to assess the reproducibility of hub-gene and model performance across independent datasets (22).
Diagnostic labels were obtained from the original GEO annotations. The diagnostic classification follows the original study annotations of each GEO dataset, which are based on the specific research diagnostic protocols defined in the original publications of their respective datasets. Disease-stage information was not consistently available across the included datasets, so the diagnostic task was defined as AD-versus-control classification (table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-1.xlsx).
Data preprocessing
For each transcriptomic dataset, probe identifiers were matched to official gene symbols using the corresponding platform annotation files. Probes mapping to the same gene were summarized by the median expression value. Expression values were log2-transformed when required (23). Within-dataset normalization was performed using the “normalizeBetweenArrays” function in the R package “limma” (24).
GSE5281 and GSE66333 were combined to form the training cohort. Batch correction was applied within this cohort using the R package sva. Principal component analysis (PCA) was used to visualize dataset-level variation before and after correction (25). External validation cohorts were processed separately. The final expression matrices were restricted to genes shared between the training cohort and each external validation cohort before external validation.
DEA
DEA was performed in the training cohort to identify genes with altered expression between AD and control samples. The analysis was conducted using the R package limma. Genes were considered differentially expressed when they met both criteria: absolute log2fold change greater than 0.585 and false discovery rate (FDR) less than 0.05. This fold change threshold corresponds approximately to a 1.5-fold expression difference and was selected to retain biologically interpretable transcriptional changes while excluding marginal technical noise, consistent with thresholds applied in published brain transcriptomic AD studies (26).
Threshold sensitivity was evaluated using alternative absolute log2fold change cutoffs combined with FDR control. Candidate-gene sets obtained under different thresholds were compared with the final prioritized genes to assess the stability of downstream gene selection. DEGs were used as the initial molecular discovery set for WGCNA.
WGCNA
WGCNA was applied to the training cohort to identify gene modules associated with AD status (27). The analysis was performed using the R package WGCNA. Sample clustering was used to inspect sample-level outliers. A soft-thresholding power was selected according to the scale-free topology fit index and mean connectivity. An adjacency matrix was constructed and transformed into a topological overlap matrix. Genes were clustered according to topological overlap dissimilarity, and co-expression modules were detected using dynamic tree cutting. Modules with highly correlated eigengenes were merged.
Each module was summarized by its module eigengene. Module eigengenes were correlated with AD status to identify disease-associated modules. For each module, the correlation coefficient, P value, gene number, and module membership distribution were recorded. The module showing the strongest association with AD was integrated with differential expression results. Other modules with nominal disease associations were reported in the supplementary materials.
Genes from the AD-associated co-expression module were intersected with DEGs. The overlapping genes represented candidates supported by altered expression and disease-related co-expression structure.
PPI network construction and candidate-gene ranking
The overlapping genes from DEA and WGCNA were mapped to the STRING protein-interaction database using the STRINGdb package. Interactions with a combined score of at least 400 were retained (28). Gene identifier conversion was performed using org.Hs.eg.db and clusterProfiler. The PPI network was constructed and analyzed using the R package igraph (29).
Five network centrality measures were calculated for each gene: degree, betweenness, closeness, eigenvector centrality, and PageRank. Genes were ranked separately according to each centrality measure. The final hub ranking was obtained by aggregating the five centrality ranks with equal weights (0.20 per metric). To assess the robustness of this weighting scheme, we performed a sensitivity analysis by (I) leaving out one centrality metric at a time and (II) randomly perturbing weights (Dirichlet distribution, 1,000 iterations). Genes that remained in the top 50 under ≥80% of perturbation runs were retained as PPI-prioritized candidates, avoiding arbitrary prior assumptions about the relative importance of individual centrality dimensions. Genes with the highest aggregate ranks were retained as PPI-prioritized candidates (30).
Genes that remained consistently prioritized under leave‑one‑centrality‑out and random weight‑perturbation analyses were carried forward to functional enrichment and diagnostic modeling.
Functional enrichment analysis
Functional enrichment analysis was performed to characterize the biological roles of PPI-prioritized genes. Gene Ontology (GO) enrichment analysis included biological process, molecular function, and cellular component categories. Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis was used to identify enriched signaling pathways. Disease Ontology (DO) analysis was used to evaluate disease-related annotations (31).
The analyses were performed using the R package clusterProfiler. Gene symbols were converted to Entrez identifiers before enrichment testing. Terms with P<0.05 and FDR <0.05 were considered significant. Enrichment results were visualized using dot plots, bar plots, chord diagrams, and enrichment networks. The enrichment results were used to interpret the biological context of the prioritized genes.
Machine-learning model development
The PPI‑prioritized genes were used as candidate predictors for AD‑versus‑control classification. Model development was restricted to the training cohort. Twelve machine‑learning algorithms were evaluated: elastic net (Enet; alpha =0.5), Lasso (alpha =1), Ridge (alpha =0), stepwise logistic regression (direction = “both”), support vector machine (SVM; radial kernel), linear discriminant analysis (LDA), glmBoost, plsRglm, random forest (RF; ntree =1,000, nodesize =5), gradient boosting machine (GBM; interaction.depth =3, shrinkage =0.001, n.trees selected by 10‑fold CV), extreme gradient boosting (XGBoost; max.depth =2, eta =1, nrounds selected by 5‑fold CV), and naïve Bayes. These algorithms were combined with four feature pre‑filtering strategies (simple all‑features, Lasso‑selected, RF‑selected, XGBoost‑selected), yielding 127 model‑feature combinations prespecified in table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-2.xlsx (32).
Hyperparameter optimization was performed exclusively within the training cohort using algorithm-specific internal CV procedures. For regularized regression models, 10-fold CV was used to select the regularization parameter (lambda) by minimizing binomial deviance. For XGBoost, 5-fold CV determined the optimal number of boosting rounds. For GBM, 10-fold CV selected the number of trees. For stepwise logistic regression, AIC-based bidirectional selection was applied within the training set. For all remaining algorithms, default parameters from the respective R packages (glmnet, e1071, MASS, mboost, plsRglm, randomForestSRC, gbm, xgboost, klaR) were used (33).
Internal CV performance [training-set CV area under the curve (AUC)] was used as the primary criterion for hyperparameter selection and for flagging models with evidence of overfitting, defined as a gap of more than 0.10 between training AUC and CV AUC. All 127 model-feature combinations and their internal performance metrics are reported in table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-3.xlsx. The three external validation cohorts (GSE110226, GSE28146, and GSE29378) were accessed only after all hyperparameters had been finalized; they served exclusively for post-hoc evaluation of generalizability and were never used for feature selection, hyperparameter tuning, model selection, or threshold calibration. The RF model, which achieved the highest internal CV AUC while showing no evidence of overfitting, was designated as the final model.
SHAP-based model interpretation
SHAP analysis was used as an interpretability method after model locking and did not replace feature selection, which had already been constrained through DEA, WGCNA, PPI prioritization, internal CV-based model development, and external validation (34). SHAP analysis was performed exclusively on the training cohort after the final model had been locked; it was not applied to any external validation cohort to prevent post-hoc data-dependent feature selection. SHAP values were calculated to estimate the contribution of each candidate gene to model predictions. Mean absolute SHAP values were used to summarize global feature importance. Bootstrap resampling was used to evaluate the stability of feature contributions.
The final candidate genes were selected by integrating mean absolute SHAP contribution (top 20% of all candidate genes), bootstrap stability [95% confidence interval (CI) of SHAP values not crossing zero after 100 resamples], PPI rank stability (retained in ≥80% of perturbation runs), and cross‑cohort diagnostic reproducibility (AUC >0.70 in all three external validation cohorts). SHAP analysis contributed to model interpretation and candidate-gene prioritization.
Diagnostic performance evaluation
Model performance was evaluated using receiver operating characteristic (ROC) analysis. AUC values were calculated using the R package pROC (35). AUC CIs were estimated using DeLong’s method or bootstrap resampling. Sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), and F1-score were calculated using the classification threshold selected in the training cohort.
Calibration was assessed using calibration curves, Brier score, and calibration slope. Decision-curve analysis was used to estimate clinical net benefit across threshold probabilities. Performance was reported separately for the training cohort, internal CV, and each external validation cohort.
Training performance, CV performance, and external validation performance across the three external cohorts (GSE110226, GSE28146, and GSE29378) were compared to assess overfitting and generalizability.
Single-gene ROC analysis was performed for the final candidate genes in the training cohort and in each external validation cohort. Genes were described as hub-genes when they showed reproducible discrimination across datasets and were supported by downstream biological-context analyses.
Transcription factor-candidate gene regulatory analysis
Transcription factor-target relationships for the final hub-genes were obtained from the TRRUST database (36). Candidate transcription factors were selected according to database confidence scores. Spearman correlation analysis was performed between transcription factor expression and candidate-gene expression in the training cohort.
A transcription factor-gene network was constructed to visualize potential upstream regulatory relationships. The analysis was used to generate regulatory hypotheses for the final hub-genes.
Immune infiltration analysis
Immune-cell infiltration was estimated using CIBERSORTx with the LM22 reference signature. The analysis was performed in the training cohort (37). Quantile normalization and 1,000 permutations were applied according to the CIBERSORTx workflow. Samples with deconvolution P<0.05 were retained for downstream immune-correlation analysis.
Spearman correlation analysis was used to examine associations between final candidate-gene expression and estimated immune-cell fractions. Correlation coefficients and adjusted P values were reported. The immune analysis was used to characterize the immunological context of the final genes in AD brain tissue.
Single-cell transcriptomic analysis
Single-cell RNA-sequencing data (GSE147047) were processed using Seurat (38). Low-quality cells were filtered according to detected gene number, UMI count, and mitochondrial transcript percentage. Data were normalized using the LogNormalize method. Highly variable genes were selected for dimensionality reduction and clustering.
PCA was performed, followed by shared nearest-neighbor graph construction and Louvain clustering. Batch correction was performed using Harmony when multiple samples or batches were included. Cell types were annotated using SingleR and checked against canonical marker genes.
Expression of the final candidate genes was visualized using feature plots, violin plots, and dot plots. The analysis was used to identify the major cellular sources of hub-genes expression.
Drug-gene enrichment and compound retrieval
Drug-gene enrichment analysis was performed after the final hub-genes had been identified (39). Drug–gene associations were obtained from DSigDB using the clusterProfiler enrichment workflow. Drug entries with P<0.05 and adjusted P<0.05 were retained.
Candidate compounds associated with enriched drug entries were retrieved from PubChem. Compound names, PubChem compound identifiers, and SMILES strings were collected. The retrieved compounds were used for drug-likeness assessment and docking-based structural prioritization.
ADMET and drug-likeness assessment
Molecular descriptors were calculated using RDKit. The descriptors included molecular weight, LogP, topological polar synthetic accessibility, hydrogen-bond donors, hydrogen-bond acceptors, and rotatable bonds (40). Drug-likeness was assessed using Lipinski, Veber, Ghose, quantitative estimate of drug-likeness, synthetic accessibility, structure-alert screening, and predicted ADMET properties.
Compounds were assigned a composite prioritization score based on the prespecified descriptor and ADMET criteria. Compounds with favorable predicted drug-likeness and ADMET profiles were retained for molecular docking (41).
Molecular docking and binding-mode analysis
Molecular docking was performed to examine the possible binding modes between prioritized compounds and selected protein structures. Protein structures were obtained from the RCSB Protein Data Bank (42). Protein and ligand structures were prepared for docking by removing nonessential molecules, adding hydrogens, assigning protonation states at physiological pH (proteins) or from SMILES strings with energy minimization (ligands), and converting both to PDBQT format.
Binding pockets were defined using co-crystallized ligand coordinates, reported functional residues, or predicted pocket locations. Grid-box centers and dimensions were recorded for each target. Docking was performed using Smina with prespecified exhaustiveness settings. When co-crystallized ligand information was available, redocking was performed, and root-mean-square deviation (RMSD) was calculated between crystallographic and redocked poses.
Docking poses were ranked by predicted binding energy (43). Representative interactions, including hydrogen bonds, hydrophobic contacts, π-π interactions, salt bridges, and halogen bonds, were visualized. Docking results were used as structural hypotheses for future pharmacological testing.
Results
Data preprocessing and DEA identified AD-associated transcriptional alterations in the discovery cohort
GSE5281 and GSE66333 were combined to establish the training cohort, comprising 91 AD samples and 78 control samples (table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-4.xlsx). PCA showed evident dataset-level separation before batch correction, indicating non-negligible inter-dataset variation between the two transcriptomic cohorts (Figure 2A). After batch correction, this dataset-driven separation was reduced, and samples from the two cohorts showed improved distributional overlap in the corrected expression space (Figure 2B).
DEA was then performed in the corrected training cohort. Using an absolute log2fold change threshold greater than 0.585 and FDR <0.05, 2,534 DEGs were identified between AD and control samples, including 1,325 upregulated and 1,209 downregulated genes in AD (Figure 2C,2D; table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-5.xlsx). These DEGs were used as the initial AD-associated transcriptional alteration set for subsequent WGCNA.
To assess whether the differential-expression screening was overly dependent on a single fold change cutoff, additional DEG sets were generated under alternative absolute log2fold change thresholds while maintaining FDR control. The number and overlap of DEGs obtained under each threshold are summarized in table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-6.xlsx. The PPI-prioritized top 50 candidate gene set showed substantial overlap (≥70%) across all tested thresholds, supporting the robustness of the transcriptomic discovery layer to fold change threshold choice.
WGCNA identified an AD-associated co-expression module that converged with differential expression signals
To determine whether the AD-associated transcriptional alterations were organized within disease-related co-expression structures, WGCNA was performed in the training cohort. A soft‑thresholding power of 6 was selected because it achieved a scale‑free topology fit index R2>0.85 while maintaining mean connectivity above 0.5 (Figure 3A-3C). Dynamic tree cutting and module merging identified multiple co-expression modules (Figure 3D).
Among these modules, the brown module showed the strongest association with AD status (R =0.62, P<0.05) and contained 1,854 genes (Figure 3E). Other modules with nominal disease associations, including the turquoise and blue modules, are reported in table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-7.xlsx. The brown module was selected because it demonstrated the strongest and most biologically interpretable association with AD status; however, additional nominally associated modules were retained in the supplementary results to preserve analytical transparency. Intersecting the brown-module genes with the 2,534 DEGs yielded 848 overlapping genes (Figure 3F). These 848 genes were therefore supported by two independent transcriptomic layers: altered expression between AD and control samples and membership in the most AD-associated co-expression module.
PPI-based rank aggregation narrowed the disease-associated gene set to 50 network-supported candidates
The 848 overlapping genes were mapped to the STRING PPI network to identify candidates occupying central positions in the AD-associated interaction structure. Five complementary network centrality metrics were calculated for each mapped gene: degree, betweenness, closeness, eigenvector centrality, and PageRank. Rank aggregation across these metrics identified 50 PPI-prioritized candidate genes for subsequent functional enrichment and diagnostic modeling (Figure 4A; table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-8.xlsx).
The 50 PPI-prioritized genes were enriched for glial differentiation, gliogenesis, ER/vesicle lumen compartments, chromatin DNA binding, and integrin binding (Figure 4B). KEGG analysis identified enriched pathways including adherens junctions, FoxO signaling, JAK-STAT signaling, Notch signaling, microRNAs in cancer, and human papillomavirus infection (Figure 4C). DO analysis showed enrichment for brain ischemia, cognitive disorder, retinal disease, and endocrine gland cancer annotations (Figure 4D). These results provided biological context for the PPI-prioritized gene set before diagnostic model construction.
Cross-cohort validation and SHAP interpretation nominated IGF1R and SPP1 as final candidate hub genes
From the 50 PPI-prioritized genes, 127 machine learning (ML) model-feature combinations were evaluated within the training cohort using algorithm-specific internal CV for hyperparameter tuning. The RF classifier achieved the highest internal CV AUC and showed no evidence of overfitting (training-CV AUC gap <0.10) and was therefore designated as the final model; its performance on external validation is reported as AUC =0.957 (95% CI: 0.922–0.984; Figure 5A). At the Youden-optimized threshold, the model showed sensitivity 0.978, specificity 0.936, PPV 0.947, NPV 0.973, F1-score 0.962, and accuracy 0.959 (table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-9.xlsx). After the final model was locked, SHAP analysis was performed exclusively on the training cohort to interpret the contribution of each gene to model predictions. SHAP values were estimated using a permutation-based approach based on the training cohort, and global feature importance was quantified as the mean absolute SHAP value across all samples. Genes were then ranked according to their mean absolute SHAP values. To assess the robustness of feature importance estimates, bootstrap resampling was performed 100 times to obtain 95% CIs for the mean absolute SHAP values. Based on the SHAP importance ranking, the top 11 genes were identified as the most influential model-contributing features for interpretation: APP, CREBBP, SUZ12, CXCR4, SNAP23, BCL6, NFKBIA, IGF1R, CEBPB, SPP1, and EZR (Figure 5B).
Single‑gene evaluation in three external validation cohorts (GSE110226, GSE28146, and GSE29378) showed that IGF1R and SPP1 both achieved AUCs >0.70 across all three cohorts (Figure 5C; tables available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-10.xlsx). By integrating SHAP contribution, bootstrap stability, PPI rank stability, and cross-cohort reproducibility, IGF1R and SPP1 were retained as the final candidate hub genes. The candidate set thus converged from 50 PPI‑network genes and 11 SHAP‑identified genes to two genes supported by network position, model contribution, and external validation.
Transcription factor and immune-infiltration analyses characterized distinct regulatory and immune associations of IGF1R and SPP1
A transcription factor-candidate gene regulatory network was constructed for the final hub genes using TRRUST-derived regulatory relationships. IGF1R related to multiple transcription factors, including TP53, SP1, POU2F1, NR3C1, KLF6, and WT1 (Figure 6A).
CIBERSORTx analysis was then used to estimate immune-cell fractions in the training cohort. Immune infiltration analysis revealed opposing association patterns for the two hub genes. IGF1R expression was negatively correlated with several B‑cell and innate immune subsets (naïve B cells, memory B cells, monocytes, γδ T cells), whereas SPP1 expression showed positive correlations with myeloid‑lineage cells (M0–M2 macrophages, neutrophils) and a negative correlation with memory B cells (Figure 6B,6C).
The combined immune-correlation network showed that IGF1R and SPP1 were associated with different immune-cell patterns rather than a uniform immune signature (Figure 6D). These results provided immune-context evidence for the two final hub genes in AD brain transcriptomic data.
Single-cell transcriptomic analysis localized IGF1R and SPP1 expression to distinct cellular contexts
Quality-control assessment showed comparable distributions of nFeature_RNA, nCount_RNA, and percent.mt between control and AD groups, indicating consistent sequencing quality (Figure 7A). PCA identified neurodevelopment- and neuronal-associated genes, including FOXG1, NEUROD6, SOX5, and PAX6, as major contributors to transcriptional heterogeneity (Figure 7B).
t-SNE analysis identified two major cellular populations, consisting primarily of neuroepithelial-like cells and neurons, with distinct distribution patterns between control and AD groups (Figure 7C). Pseudotime trajectory analysis further demonstrated altered trajectory distributions in AD samples compared with controls (Figure 7D).
Feature expression analysis showed that IGF1R was predominantly expressed in neuroepithelial-like cells and was significantly reduced in AD relative to controls (Wilcoxon test, P<0.01), whereas SPP1 expression remained sparse and showed no significant group difference (Figure 7E).
Drug-gene enrichment and ADMET filtering prioritized compounds linked to the final hub gene signature
Because IGF1R and SPP1 were the only genes that met all prespecified selection criteria—network centrality, SHAP contribution, bootstrap stability, and cross‑cohort AUC >0.70 across all three external validation cohorts—they were designated as the final hub genes and used as the query set for drug–gene enrichment to generate exploratory, hypothesis‑driven compound‑prioritization leads rather than validated therapeutic candidates.
Drug–gene enrichment analysis identified 176 significantly enriched drug signatures associated with the final hub-gene set (P<0.05 and adjusted P<0.05; Figure 8A; table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-11.xlsx). PubChem-based compound retrieval yielded 445 related small molecules for downstream drug-likeness and ADMET evaluation.
Molecular descriptor and ADMET profiling classified the 445 compounds into five prioritization grades. Thirty-seven compounds were assigned grade A according to the prespecified composite drug-likeness and ADMET criteria and were retained for molecular docking (Figure 8B; table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-12.xlsx). This filtering step reduced the initial drug-gene enrichment output to a smaller set of compounds with more favorable predicted physicochemical and pharmacokinetic properties.
Molecular docking identified four top-ranked compound-target binding poses
Molecular docking was performed between the 37 grade-A compounds and the selected target protein structures 3D94 and 5FXS. Docking was conducted at pH 7.4 with exhaustiveness set to 16, and binding pockets were defined according to the prespecified pocket-definition strategy (table available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-13.xlsx). Molecular docking was performed at pH 7.4±0.2 to mimic physiological conditions. Redocking of co-crystallized ligands produced RMSD values of 2.3 Å for 3D94 and 1.9 Å for 5FXS, confirming the reproducibility of the docking protocol (tables available at https://cdn.amegroups.cn/static/public/JMAI-2026-1-0025-14.xlsx).
Four representative compounds showed predicted binding energies better than −9.0 kcal/mol (Figure 9). The best-ranked ligand, C1CN(CCC12C(=O)NCN2C3=CC=CC=C3)CCCC(=O)C4=CC=C(C=C4)F, docked with 3D94 with a predicted binding energy of −10.5 kcal/mol. The binding pose showed hydrophobic contacts with LEU975, VAL983, ALA1021, and PHE1124, as well as hydrogen-bond interactions involving CYS1111, VAL1113, ASP1123, and PHE1124, with interaction distances ranging from approximately 2.2 to 3.8 Å (Figure 9A).
The compound CN1CCN(CC1)C2=NC3=C(C=CC(=C3)Cl)NC4=CC=CC=C42 docked with 5FXS with a predicted binding energy of −9.2 kcal/mol. Its binding pose included hydrophobic contacts involving LEU1005 and ALA1031, hydrogen-bond interactions with ASP1086, and a halogen-bond interaction involving ASP1153 (Figure 9B).
Two additional compounds, CC1(CCC(C2=C1C=CC(=C2)NC(=O)C3=CC=C(C=C3)C(=O)O)(C)C)C and COC1=C2C3=C(C(=O)CC3)C(=O)OC2=C4C5C=COC5OC4=C1, docked with 5FXS with predicted binding energies of −9.1 kcal/mol. Their binding poses were mainly stabilized by hydrophobic contacts involving LEU1005, ALA1031, VAL1013, LYS1033, MET1079, and MET1082, with additional hydrogen-bond interactions observed in selected poses (Figure 9C,9D).
Together, these docking results ranked four compounds as the leading structure-based candidates derived from the hub-gene-guided compound-prioritization workflow.
Discussion
In this study, we established a leakage-free transcriptomic secondary-analysis framework integrating DEA, WGCNA, PPI-based network prioritization, ensemble ML, SHAP interpretation, immune deconvolution, single-cell transcriptomic analysis, and exploratory compound prioritization to investigate molecular alterations associated with AD. Through this multistage workflow, IGF1R and SPP1 emerged as the final candidate hub genes supported simultaneously by differential expression, co‑expression structure, network centrality, machine‑learning contribution, biological interpretability, and external cross‑cohort reproducibility. Compared with studies relying on a single analytical strategy, the present framework attempted to reduce false‑positive prioritization by requiring candidate genes to remain consistently supported across multiple independent analytical dimensions.
The transcriptomic and network‑level findings further support the concept that AD is not solely a neuron‑centered degenerative disorder but a complex systems‑level disease involving extensive glial, immune, metabolic, and vascular dysregulation. DEA identified widespread transcriptional alterations in AD brain tissue, while WGCNA demonstrated that many of these dysregulated genes were concentrated within disease‑associated co‑expression structures. Functional enrichment analyses consistently highlighted pathways related to glial differentiation, gliogenesis, JAK-STAT signaling, Notch signaling, and FoxO signaling, together with disease annotations involving cognitive dysfunction and cerebral ischemia (44,45). The JAK/STAT pathway has been recognized as a critical driver of neuroinflammation in AD, mediating both innate and adaptive immune responses in the brain (46). These findings are consistent with increasing evidence that neuronal injury and disease progression in AD are collectively contributed by sustained neuroinflammation, glial activation, vascular dysfunction, and compromised metabolic signaling. Rather than representing isolated molecular events, the observed transcriptional alterations appear to form interconnected regulatory networks centered on the neuronal-glial-immune axis (47). The convergence of differential expression, co‑expression topology, and PPI‑network structure suggests that genes occupying central positions within these interconnected pathways may represent biologically meaningful disease‑associated candidates.
An important methodological aspect of this study is the emphasis on leakage-free model development and external reproducibility. Transcriptomic machine-learning studies are frequently affected by optimistic bias caused by feature leakage, inappropriate preprocessing across datasets, or insufficient separation between model development and validation procedures. To minimize these risks, feature discovery, network construction, hyperparameter tuning, threshold selection, SHAP interpretation, and model selection were restricted to the training cohort, while GSE110226, GSE28146, and GSE29378 were reserved exclusively for external validation. Under this framework, the RF model achieved stable discrimination performance, with an AUC approaching 0.96 in training-cohort internal CV and external validation analyses, suggesting that the prioritized gene signature retained reproducible diagnostic information despite inter-cohort heterogeneity. AD exhibits substantial biological and technical variability across populations and transcriptomic platforms; therefore, the preservation of model performance across external validation cohorts supports the robustness of the analytical workflow, while still requiring prospective clinical validation.
IGF1R and SPP1 showed the most reproducible performance across external cohorts and were selected as the final hub genes. Both maintained AUC >0.70 in all validation cohorts, retaining single‑gene discriminative value. These findings do not establish clinically validated biomarkers, but they suggest that both genes may represent biologically informative molecular indicators associated with AD pathology. SHAP analysis further demonstrated that IGF1R contributed strongly and consistently to model predictions, indicating that its expression pattern may reflect central disease‑related regulatory processes rather than cohort‑specific variation.
From a biological perspective, the two hub genes appear to represent partially distinct yet complementary pathological dimensions of AD. IGF1R is tightly linked to insulin/IGF signaling, neuronal survival, metabolic regulation, oxidative stress response and synaptic maintenance, all of which have been implicated in brain insulin resistance and neurodegenerative progression in AD (3-5,48). The IGF signaling pathway is crucial for CNS development and metabolism, and its dysregulation—particularly disruptions in the downstream PI3K/AKT and MAPK cascades—is known to increase AD susceptibility. These pathways regulate neuronal survival, tau phosphorylation, and synaptic plasticity. Dysregulated IGF1R signaling has also been linked to increased amyloid‑β production through altered APP processing and reduced α‑secretase activity, suggesting a broader role in AD core pathology. Dysregulated insulin signaling contributes to impaired neuronal plasticity, tau hyperphosphorylation, mitochondrial dysfunction, and synaptic degeneration. In the present study, the transcription factor (TF)-gene regulatory analysis further indicated that IGF1R was connected to multiple high‑confidence transcription factors, including TP53, SP1, POU2F1, and NR3C1. This suggests that IGF1R may function as an integrative signaling node linking stress‑response, inflammatory, and metabolic pathways, which may partly explain why IGF1R consistently exhibited strong SHAP contributions in the machine‑learning models.
SPP1 was more strongly associated with inflammatory and immune mechanisms. SPP1 (osteopontin) has been implicated in microglial activation, leukocyte recruitment, extracellular‑matrix remodeling, and chronic inflammatory amplification in neurodegenerative microenvironments (49,50). Recent evidence has shown that cortical SPP1 expression is associated with faster cognitive decline and greater odds of common neuropathologies in older adults, and that osteopontin interacts directly with APP and MAPT, activating the NF‑κB pathway and amplifying neuroinflammation (51). Consistent with this, a 2023 review contextualizes osteopontin as an inflammatory cytokine and biomarker of AD, central to microglial activation and neuroinflammation (49). Immune‑infiltration analysis showed that SPP1 expression positively correlated with macrophage and neutrophil signatures, whereas IGF1R negatively correlated with several adaptive and innate immune‑cell subsets. These findings suggest that the two genes may reflect distinct immunobiological states within AD tissue—SPP1 more strongly linked to inflammatory myeloid activation and IGF1R more closely associated with neuronal‑metabolic regulation. The observed immune‑correlation patterns are consistent with the current understanding that innate immune activation and chronic microglial dysregulation play central roles in AD progression, as also demonstrated by recent integrative bioinformatics studies that employed CIBERSORT to construct immune infiltration‑related diagnostic models for AD (52).
Single-cell transcriptomic analysis further strengthened the biological interpretability of the candidate genes by clarifying their cell-type-specific expression contexts. IGF1R expression was enriched in neuroepithelial-like and neuronal subpopulations and varied along pseudotime trajectories, linking this gene to neuronal differentiation, survival, and plasticity. In contrast, SPP1 expression was sparse and appeared mainly in inflammation-associated cellular contexts, supporting its potential relationship with immune activation. These findings indicate that IGF1R and SPP1 may reflect distinct cellular dimensions of AD-associated molecular dysregulation: IGF1R is more closely related to neuronal and neuroepithelial-like cell states, whereas SPP1 is more closely associated with inflammatory cellular states. Recent single-cell transcriptomic studies have shown that AD-associated transcriptional dysregulation differs across neuronal and non-neuronal cell types, highlighting the importance of cell-type-specific context in interpreting disease-related molecular signatures. Thus, AD-associated molecular dysregulation should be interpreted as a cell-type-dependent process rather than as a uniform whole-tissue signal. Future studies targeting these pathways should consider cell-type-specific and microenvironmental heterogeneity when evaluating the biological and translational relevance of IGF1R and SPP1.
Beyond biomarker prioritization, this study also explored a preliminary translational extension through drug-gene enrichment, ADMET‑based filtering, and molecular docking. After integrating transcriptomic evidence and network prioritization, 37 compounds with favorable predicted drug‑likeness and ADMET profiles were retained for docking analysis. Several compounds demonstrated predicted binding energies below −9.0 kcal/mol, including one ligand showing a binding energy of −10.5 kcal/mol against the selected protein structure. Structural interaction analyses suggested that hydrophobic interactions, hydrogen bonding, and halogen‑bond stabilization may contribute to the predicted binding affinity and specificity of these compounds. Similar computational approaches combining virtual screening, ADMET prediction, and molecular docking have been widely used in AD drug discovery, identifying natural product-derived compounds with favorable pharmacokinetic profiles (53). These docking results remain computational hypotheses rather than validated therapeutic findings, but they provide a rational starting point for future medicinal‑chemistry optimization and experimental validation.
The translational value of this study lies not only in identifying DEGs but also in establishing a biologically interpretable analytical chain linking transcriptomic dysregulation, co-expression structure, network centrality, machine-learning discrimination, immune context, single-cell transcriptomic evidence, and structure-based compound prioritization. This multilevel strategy provides a more coherent framework connecting molecular discovery with downstream hypothesis generation for diagnostic development and therapeutic exploration (54). By integrating independent evidence across multiple analytical layers, the study attempted to improve both the interpretability and reproducibility of candidate‑gene prioritization in AD.
Several limitations should be acknowledged. This study was based entirely on retrospective secondary analyses of publicly available transcriptomic datasets, and the findings therefore remain dependent on dataset quality, cohort composition, and platform characteristics. Batch correction and external validation were performed, but residual technical heterogeneity may still influence gene‑expression estimates. Residual confounding from unobserved clinical variables—including medication use, comorbid conditions, and postmortem interval—cannot be excluded and may independently influence gene expression estimates. Individual external validation cohorts contained modest sample sizes, which limits the precision of AUC point estimates and reduces statistical power to detect performance differences across disease stages. Disease‑stage information and detailed clinical phenotypes were not consistently available across all cohorts, limiting the ability to evaluate stage‑specific molecular dynamics. The machine‑learning framework focused on transcriptomic signatures and did not incorporate multimodal clinical variables, imaging markers, or proteomic information that may improve diagnostic performance. Immune deconvolution and single‑cell analyses were computationally inferred and require orthogonal biological validation. The compound‑prioritization and molecular‑docking analyses were exploratory in nature; docking scores alone cannot predict pharmacological efficacy, blood‑brain barrier penetration, toxicity, or in vivo therapeutic activity. No experimental validation [e.g., quantitative polymerase chain reaction (qPCR), immunohistochemistry, western blotting, CRISPR perturbation, or animal model testing] was performed for the identified hub genes or the prioritized compounds. Therefore, IGF1R and SPP1 should be interpreted as computationally nominated candidates, not as validated diagnostic or therapeutic targets. All docking-based binding predictions require orthogonal biochemical and cellular assays.
Future studies should focus on prospective multicenter validation, integration of multi‑omics and clinical data, mechanistic characterization of IGF1R‑ and SPP1‑associated pathways, and experimental validation in cellular and animal models. Additional efforts are needed to evaluate whether these candidate genes can be detected reproducibly in accessible biospecimens such as blood or CSF and whether they retain diagnostic value across different disease stages and comorbid conditions (55). From a therapeutic perspective, future work should combine molecular dynamics simulation, target‑binding assays, functional perturbation experiments, and brain‑targeted drug‑delivery strategies to evaluate the translational feasibility of the prioritized compounds. Together, these may contribute to the development of more biologically informed diagnostic and therapeutic strategies for AD (56).
Conclusions
This study established a leakage-free transcriptomic framework integrating DEA, WGCNA, PPI prioritization, and ensemble ML to identify IGF1R and SPP1 as reproducible AD-associated candidate hub genes, supported by network centrality, SHAP contribution, and cross-cohort validation. Exploratory molecular docking further nominated four compound leads for future pharmacological investigation. All findings represent computational hypotheses requiring experimental and clinical validation.
The authors declare that no artificial intelligence tools were used in the preparation, writing, analysis, or revision of this manuscript.
Footnote
Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0025/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-1-0025/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was conducted as a secondary analysis of publicly available, de-identified transcriptomic datasets. No new human subjects research, animal experiments, clinical specimens, or patient-identifiable data were generated or collected; therefore, additional institutional ethics approval and informed consent were not required.
References
- 2023 Alzheimer's disease facts and figures. Alzheimers Dement 2023;19:1598-695. [Crossref] [PubMed]
- Jack CR Jr, Andrews JS, Beach TG, et al. Revised criteria for diagnosis and staging of Alzheimer's disease: Alzheimer's Association Workgroup. Alzheimers Dement 2024;20:5143-69. [Crossref] [PubMed]
- Zhang J, Zhang Y, Wang J, et al. Recent advances in Alzheimer's disease: Mechanisms, clinical trials and new drug development strategies. Signal Transduct Target Ther 2024;9:211. [Crossref] [PubMed]
- Zheng Q, Wang X. Alzheimer's disease: insights into pathology, molecular mechanisms, and therapy. Protein Cell 2025;16:83-120. [Crossref] [PubMed]
- Heneka MT, van der Flier WM, Jessen F, et al. Neuroinflammation in Alzheimer disease. Nat Rev Immunol 2025;25:321-52. [Crossref] [PubMed]
- Palmqvist S, Whitson HE, Allen LA, et al. Alzheimer's Association Clinical Practice Guideline on the use of blood-based biomarkers in the diagnostic workup of suspected Alzheimer's disease within specialized care settings. Alzheimers Dement 2025;21:e70535. [Crossref] [PubMed]
- Mielke MM, Fowler NR. Alzheimer disease blood biomarkers: considerations for population-level use. Nat Rev Neurol 2024;20:495-504. [Crossref] [PubMed]
- Kodam P, Sai Swaroop R, Pradhan SS, et al. Integrated multi-omics analysis of Alzheimer's disease shows molecular signatures associated with disease progression and potential therapeutic targets. Sci Rep 2023;13:3695. [Crossref] [PubMed]
- Hampel H, Hu Y, Cummings J, et al. Blood-based biomarkers for Alzheimer's disease: Current state and future use in a transformed global healthcare landscape. Neuron 2023;111:2781-99. [Crossref] [PubMed]
- Nowell J, Crook H, de Leon MJ, et al. Advances in the drug treatment of Alzheimer's disease: pathophysiology and mechanisms of action. BMJ 2026;393:78881. [Crossref] [PubMed]
- Badhwar A, McFall GP, Sapkota S, et al. A multiomics approach to heterogeneity in Alzheimer's disease: focused review and roadmap. Brain 2020;143:1315-31. [Crossref] [PubMed]
- Wirka RC, Pjanic M, Quertermous T. Advances in Transcriptomics: Investigating Cardiovascular Disease at Unprecedented Resolution. Circ Res 2018;122:1200-20. [Crossref] [PubMed]
- Brunwasser SM, Warner AK, Rosas-Salazar C, et al. Advancing birth cohort studies using administrative and other research-independent data repositories: Opportunities and challenges. J Allergy Clin Immunol 2025;155:1727-35. [Crossref] [PubMed]
- Gupta A, Chug A, Singh AP. Feature extraction in sensor plant disease datasets using reformed membership functions independent of class variables. Sci Rep 2026;16:4115. [Crossref] [PubMed]
- Lv Y, Fan G. Network-Driven Insights into Plant Immunity: Integrating Transcriptomic and Proteomic Approaches in Plant-Pathogen Interactions. Int J Mol Sci 2026;27:1242. [Crossref] [PubMed]
- Xiao W, Yuan L, Ran T, et al. CSPER: Curiosity-driven Safety-Prioritized Experience Replay for Mapless Navigation using Deep Reinforcement Learning. IEEE Trans Cogn Dev Syst 2026;18:1028-42.
- Singh S, Kumar M, Sengar V, et al. Ensemble learning for air quality index prediction: integrating gradient boosting, XGBoost, and stacking with SHAP-based interpretability. Sci Rep 2026;16:8544. [Crossref] [PubMed]
- Hu C, Guan M, Mi F, et al. Multimodal SERS Biosensing Platforms: Emerging Opportunities for Ultrasensitive Biomarker Detection and Intelligent Diagnostics. ACS Sens 2026;11:1794-830. [Crossref] [PubMed]
- Shen Y, Zhang P, Luo J, et al. Artificial Intelligence Drives Advances in Multi-Omics Analysis and Precision Medicine for Sepsis. Biomedicines 2026;14:261. [Crossref] [PubMed]
- Jiménez-Gracia L, Maspero D, Aguilar-Fernández S, et al. Interpretable inflammation landscape of circulating immune cells. Nat Med 2026;32:633-44. [Crossref] [PubMed]
- Brown GS, Wengler J, Fabelico AJS, et al. Using semantic search to find publicly available gene-expression datasets. Bioinformatics 2026;42:btag053. [Crossref] [PubMed]
- Dong Z, Meng X, Wang L. Ion-Channel-Mediated Drug Repurposing Opportunities Validated by Single-Cell Perturbation in Colorectal Cancer. Int J Mol Sci 2026;27:3412. [Crossref] [PubMed]
- Shahin H, Seylani A, Mahalwar G. 26-A-20734-ACC from bone marrow to heart: how long non-coding RNAs in multiple myeloma influence cardiovascular risk. J Am Coll Cardiol 2026;87:A1120.
- Jin J, Jin Y, Lu B, et al. Identification and histological validation of autophagy-related core genes ADRB2 and PLK2 in keloids, with integrated immune infiltration analysis. Front Immunol 2026;17:1724230. [Crossref] [PubMed]
- Wang Z, Hou J, Chen Y, et al. Integrated Multi-Omics and Machine Learning Framework Identifies Diagnostic Signatures and Druggable Targets in Breast Cancer. Genes (Basel) 2026;17:396. [Crossref] [PubMed]
- Wang Z, Li J, Lei B. AI-driven discovery of EZH2 as a core target and potential inhibitors for breast cancer. J Med Artif Intell 2026;9:32.
- Duan S, Zhou Q, Shi Y, et al. Immune Infiltration and Mitochondrial Function in Diabetic Kidney Disease: WGCNA and Machine Learning Identified Hub Genes with Clinical Validation. Int J Mol Sci 2026;27:4696. [Crossref] [PubMed]
- Bathla D, Vindal V. Protein interaction network analysis: STRING and CYTOSCAPE. In: Vemireddy LNR, Reddyyamini B, Amaravathi Y, editors. Harnessing Genomic Tools for Crop Improvement. Cambridge: Academic Press; 2026:393-418.
- Wang C, Deng Y, Yang R. Integrated network toxicology and immune profiling identify ESR1 as a potential hub linking Benzo[a]pyrene exposure to gastric cancer risk. Discov Oncol 2026;17:746. [Crossref] [PubMed]
- Ikrar T, Siahaan SC, Hendarto H, et al. Multi-Target Modulation of Metabolic and Steroidogenic Pathways by <i>Cinnamomum burmannii</i> and <i>Myristica fragrans</i> in Polycystic Ovary Syndrome: An Integrative Transcriptomics, Metabolomic, Pharmacoinformatics and Experimental Validation. Nutrients 2026;18:1305. [Crossref] [PubMed]
- Parrotta EI, Benedetto GL, Cuda G, et al. Proteomic biomarkers in rotator cuff disease: Current evidence and translational implications from a systematic review and network analysis. J Orthop Translat 2026;57:101069. [Crossref] [PubMed]
- Teodorescu V, Obreja Brașoveanu L. Assessing the validity of k-fold cross-validation for model selection: Evidence from bankruptcy prediction using random forest and XGBoost. Computation 2025;13:127.
- Imani M, Beikmohammadi A, Arabnia HR. Comprehensive analysis of random forest and XGBoost performance with SMOTE, ADASYN, and GNUS under varying imbalance levels. Technologies 2025;13:88.
- Wang L, Wu M. Research on bearing fault diagnosis based on machine learning and SHAP interpretability analysis. Sci Rep 2025;15:41242. [Crossref] [PubMed]
- Riesthuis P, Otgaar H. On the use of receiver operating characteristic area under the curve in eyewitness memory research. Legal and Criminological Psychology 2025;30:212-30.
- Han H, Cho JW, Lee S, et al. TRRUST v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic Acids Res 2018;46:D380-6. [Crossref] [PubMed]
- Chen B, Khodadoust MS, Liu CL, et al. Profiling Tumor Infiltrating Immune Cells with CIBERSORT. Methods Mol Biol 2018;1711:243-59. [Crossref] [PubMed]
- Kang Y. Comparison of Different in Vitro Models of Alzheimer's Disease Using Re-Analysis of Scrna-Seq Data. In: Proceedings of the 2022 International Conference on Intelligent Medicine and Health. Xiamen: Association for Computing Machinery; 2022:84-90.
- Zhang S, Xiang X, Liu L, et al. Bioinformatics Analysis of Hub Genes and Potential Therapeutic Agents Associated with Gastric Cancer. Cancer Manag Res 2021;13:8929-51. [Crossref] [PubMed]
- Aires-de-Sousa J. GUIDEMOL: A Python graphical user interface for molecular descriptors based on RDKit. Mol Inform 2024;43:e202300190. [Crossref] [PubMed]
- Qian L, Wang ZF, He L, et al. The Mechanism of Buqi Jianzhong Decoction in Treating Hepatocellular Carcinoma via AI-assisted Network Pharmacology and Experimentation. Curr Comput Aided Drug Des 2026; [Crossref] [PubMed]
- Goodsell DS, Zardecki C, Di Costanzo L, et al. RCSB Protein Data Bank: Enabling biomedical research and drug discovery. Protein Sci 2020;29:52-65. [Crossref] [PubMed]
- Ramírez D, Caballero J. Is It Reliable to Use Common Molecular Docking Methods for Comparing the Binding Affinities of Enantiomer Pairs for Their Protein Target? Int J Mol Sci 2016;17:525. [Crossref] [PubMed]
- Song T, Zhang Y, Zhu L, et al. The role of JAK/STAT signaling pathway in cerebral ischemia-reperfusion injury and the therapeutic effect of traditional Chinese medicine: A narrative review. Medicine (Baltimore) 2023;102:e35890. [Crossref] [PubMed]
- Wu JW, Wang BX, Shen LP, et al. Investigating the Potential Therapeutic Targeting of the JAK-STAT Pathway in Cerebrovascular Diseases: Opportunities and Challenges. Mol Neurobiol 2025;62:9338-64. [Crossref] [PubMed]
- Rusek M, Smith J, El-Khatib K, et al. The Role of the JAK/STAT Signaling Pathway in the Pathogenesis of Alzheimer's Disease: New Potential Treatment Target. Int J Mol Sci 2023;24:864. [Crossref] [PubMed]
- Maiese K. The Metabolic Basis for Nervous System Dysfunction in Alzheimer's Disease, Parkinson's Disease, and Huntington's Disease. Curr Neurovasc Res 2023;20:314-33. [Crossref] [PubMed]
- Miao J, Zhang Y, Su C, et al. Insulin-Like Growth Factor Signaling in Alzheimer's Disease: Pathophysiology and Therapeutic Strategies. Mol Neurobiol 2025;62:3195-225. [Crossref] [PubMed]
- De Schepper S, Ge JZ, Sierksma A, et al. Perivascular cells induce microglial phagocytic states and synaptic engulfment via SPP1 in mouse models of Alzheimer's disease. Nat Neurosci 2023;26:406-15. [Crossref] [PubMed]
- Argandona Lopez C, Brown AM. Microglial- neuronal crosstalk in chronic viral infection through mTOR, SPP1/OPN and inflammasome pathway signaling. Front Immunol 2024;15:1368465. [Crossref] [PubMed]
- Lopes KP, Yu L, Shen X, et al. Associations of cortical SPP1 and ITGAX with cognition and common neuropathologies in older adults. Alzheimers Dement 2024;20:525-37. [Crossref] [PubMed]
- Xiao Z, Hu R, Liu WL, et al. Identification and immunological characterization of genes associated with ferroptosis in Alzheimer's disease and experimental demonstration. Mol Med Rep 2024;30:155. [Crossref] [PubMed]
- Gheidari D, Mehrdad M, Karimelahi Z. Virtual screening, ADMET prediction, molecular docking, and dynamic simulation studies of natural products as BACE1 inhibitors for the management of Alzheimer's disease. Sci Rep 2024;14:26431. [Crossref] [PubMed]
- Alshabrmi FM, Aba Alkhayl FF, Rehman A. Novel drug discovery: Advancing Alzheimer's therapy through machine learning and network pharmacology. Eur J Pharmacol 2024;976:176661. [Crossref] [PubMed]
- Liu Q, Ling J, Li Z, et al. Advances in lymphoma biomarkers research based on proteomics technology (Review) Oncol Rep 2025;54:108. [Crossref] [PubMed]
- Günaydın T, Varlı S. Computer-Aided Decision Support Systems of Alzheimer's Disease Diagnosis - A Systematic Review. Curr Med Imaging 2025;21:e15734056359358. [Crossref] [PubMed]
Cite this article as: Wang Z, Hou J, Zhu Y, Guan C, Vengusamy S. Multi-cohort transcriptomic analysis and ensemble machine learning identify IGF1R and SPP1 as candidate diagnostic genes and nominate compound leads for Alzheimer’s disease. J Med Artif Intell 2026;09:68.

