Comparison of machine-learning-based auto-segmentation to manual segmentation of knee and shoulder CT scans
Original Article

Comparison of machine-learning-based auto-segmentation to manual segmentation of knee and shoulder CT scans

Abdulganeey Olawin, Tom Gale, Emily C. Gray, Clarissa LeVasseur, Raghav Ramraj, Emelia Krakora, Albert Lin, Kenneth Urish, Gele Moloney, MaCalus Hogan, William Anderst

1Department of Orthopaedic Surgery, University of Pittsburgh, Pittsburgh, PA, USA; 2Foot and Ankle Injury Research [F.A.I.R] Group, University of Pittsburgh, Pittsburgh, PA, USA

Contributions: (I) Conception and design: EC Gray, T Gale, W Anderst; (II) Administrative support: A Olawin, EC Gray; (III) Provision of study materials or patients: G Moloney, A Lin, K Urish, M Hogan, W Anderst; (IV) Collection and assembly of data: A Olawin, EC Gray, T Gale, R Ramraj, E Krakora; (V) Data analysis and interpretation: A Olawin; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: William Anderst, PhD. Department of Orthopaedic Surgery, University of Pittsburgh, 3820 S Water Street, Pittsburgh, PA 15203, USA. anderst@pitt.edu

Background: Manual segmentation is the gold standard for segmenting bone tissue in computed tomography (CT) scans. Recently developed auto-segmentation algorithms have the potential for substantial time saving. However, most segmentation algorithms lack external validation and have not been evaluated for their robustness to pathology and age-associated bony changes. This study aimed to evaluate the accuracy of a commercially available auto-segmentation algorithm across different joints in various patient cohorts. We hypothesized that (I) auto-segmentation accuracy would not differ significantly among the individual bones composing a given joint, and (II) auto-segmentation would perform more accurately in younger, healthy subjects without known musculoskeletal pathology than in older subjects or those with known musculoskeletal pathology.

Methods: This study evaluated the accuracy of a commercially available auto-segmentation software for CT scans of the knee and shoulder using manual segmentation as the reference standard for 136 knees and 59 shoulders. Scans from healthy individuals and those with musculoskeletal pathologies were included. Accuracy was evaluated through Dice similarity coefficient (DSC), 1-pixel outer edge DSC, and average surface deviation (ASD) of 3D models created from the segmented bones.

Results: Auto-segmentation yielded predominantly high DSC scores for whole-bone masks (average DSC ranging from 0.905 to 0.983), although a few cases showed relatively low performance. Edge accuracy was less accurate, with average 1-pixel edge ranging from 0.376 to 0.654. Mean ASD values were within approximately 1 mm for knees and shoulders, although knee ASD differed among groups (P<0.001). The evaluated commercial auto-segmentation accuracy was found to be dependent upon bone, patient age, and the presence of pathology. Specifically, tibial DSC exceeded femoral DSC in the knee (0.967±0.045 vs. 0.933±0.129; P<0.001), humeral DSC exceeded scapula DSC in the shoulder (0.979±0.029 vs. 0.957±0.022; P<0.001), younger healthy knees had lower DSC than older healthy and older pathologic knees (P<0.001 for both), older pathologic shoulders had lower DSC than younger healthy shoulders (P<0.001), and knee ASD was greater in older healthy knees than younger healthy and younger pathologic knees (P<0.001 and P=0.02, respectively). Auto-segmentation was considerably faster than manual segmentation; approximately 1 minute automated versus 30 to 60 minutes manual.

Conclusions: The evaluated commercial auto-segmentation algorithm was rapid and demonstrated high whole-bone volumetric accuracy for CT-based knee and shoulder bone segmentation. However, boundary-level accuracy was comparatively poor. It’s possible these limitations may be overcome by developing auto-segmentation algorithms trained specifically for patients based on their age and pathology.

Keywords: Auto-segmentation; computed tomography (CT); Dice similarity coefficient (DSC)


Received: 12 March 2026; Accepted: 10 July 2026; Published online: 26 August 2026.

doi: 10.21037/jmai-2026-0059


Highlight box

Key findings

• This study evaluated the accuracy of a commercially available auto-segmentation algorithm (Simpleware) against manual segmentation as the reference standard across 136 knees and 59 shoulders, spanning healthy and pathologic cohorts. Whole-bone volumetric accuracy was generally high, with mean Dice similarity coefficient (DSC) values ranging from 0.905 to 0.983 for knee bones and 0.957 to 0.979 for shoulder bones. However, boundary-level accuracy was substantially lower, with 1-pixel DSC values ranging from 0.376 to 0.654 across all groups. Segmentation accuracy was influenced by bone type, patient age, and pathology in a joint-specific manner. Auto-segmentation was approximately 36- to 50-fold faster than manual segmentation.

What is known and what is new?

• Manual segmentation of computed tomography scans is the accepted gold standard for bone model generation but is time-consuming and subject to interobserver variability. Prior studies have reported high DSC values for automated knee and shoulder segmentation; however, none have evaluated boundary-level accuracy using a 1-pixel DSC metric, nor have they systematically stratified results by age and pathology.

• This study introduces 1-pixel DSC as a complementary validation metric and demonstrates that volumetric accuracy does not reliably reflect cortical edge precision.

What is the implication, and what should change now?

• Commercially available auto-segmentation algorithms may be sufficient for applications requiring gross volumetric bone models but should be used cautiously for high-precision tasks such as implant fitting, fracture characterization, or 2D-to-3D registration where boundary accuracy is critical. Future algorithm development should incorporate age- and pathology-stratified training datasets, and validation studies should routinely report boundary-level metrics alongside conventional DSC.


Introduction

Background

The ability to accurately and quickly segment bone models from computed tomography (CT) is of high importance in a clinical setting. These bone models can be used to improve the diagnosis of bone-related diseases, generate personalized treatment plans, and effectively plan surgical procedures. For example, accurate three-dimensional bone models are needed for fracture characterization and reduction planning to reduce malunion and nonunion rates (1-3), for precision-dependent applications such as total knee arthroplasty (TKA) (1), and for biomechanical simulations to estimate failure loads (4). Manual segmentation, which requires a technician to manually identify bone tissue slice-by-slice in the CT, is the gold standard in image segmentation (Figure 1A) despite being labor-intensive and time-consuming (1,2). With auto-segmentation, the user imports the CT scan and the computer uses a trained algorithm to generate a bone model in a matter of seconds, compared to over an hour required for some manual segmentation methods (5). Auto-segmentation offers several additional advantages, including removal of interobserver variability and potential improvements in accuracy (5-7). This approach may reduce the time from diagnosis to treatment and, in turn, decrease healthcare costs.

Figure 1 Overview of the segmentation comparison workflow. (A) A representative CT scan with a manually segmented femur, (B) a single slice of the manual segmentation overlaid on auto-segmentation with gray area indicating matched pixels, blue indicating a manually segmented mask (manual) only and orange indicating an auto-segmented mask (auto), (C) a 1-pixel edge mask, (D) DSC calculation for entire mask vs. 1-pixel DSC for the edges. CT, computed tomography; DSC, Dice similarity coefficient.

Rationale and knowledge gap

Metrics such as the Dice similarity coefficient (DSC) and the average surface distance/deviation have been used to evaluate the accuracy of deep learning models in generating correct segmentation masks (5). The DSC computes the spatial overlap between the ground truth mask (manually segmented mask) and the predicted mask (auto-segmented mask) (8). While there is no universal DSC threshold that defines musculoskeletal clinical acceptability, values greater than 0.90 are typically considered as high volumetric overlap (9-11). However, the overall DSC may be insufficient for evaluating the accuracy in segmenting the outer edges of bones where a fracture may occur or where osteophytes may develop with age because bone edges comprise only a small portion of the overall bone volume.

Accuracy of outer surfaces may be clinically important in cases involving abnormal morphologic features (e.g., osteophytes, Hill-Sachs legions) (12) or when the masks are used for high precision tasks such as 2D to 3D bone registration (13). While several studies have evaluated the accuracy of auto-segmented models using the DSC and average surface deviation (ASD), no studies have calculated the accuracy of the isolated edge of the auto-segmented masks using DSC (1-pixel DSC) (6,9,14). Poor cortical-edge segmentation may alter estimated fracture morphology, osteophyte size, implant-fit calculations, joint-contact mechanics, and 2D-to-3D registration accuracy, which may in turn affect clinical decision-making and treatment choice. Therefore, a boundary-focused metric such as 1-pixel DSC can provide information that is not captured by whole-mask DSC alone.

A few studies have assessed automated segmentation accuracy. For knee segmentation, one study of 50 CT scans from asymptomatic subjects reported DSC values as high as 0.98 (6), while another knee-based study using CT scans of three cadavers reported an ASD ranging from −0.06 to 1.40 mm (14). In contrast, a study of shoulders reported a DSC of approximately 0.97 and an average error of 1.0–2.55 mm based on scans from 60 normal and 56 osteoarthritic scapulae (9). However, none of these studies stratified their results by age and pathology. It is important to evaluate the accuracy of machine learning segmentation in older and pathologic joints because age and pathology are associated with abnormal bone surface features such as osteophytes, reduced bone mineral density, and joint space narrowing (15). Additionally, automated segmentation accuracy may not be consistent across the different bones within a joint, given that bone shape, density, and resulting CT voxel values can vary considerably from bone to bone (16).

Auto-segmentation using machine learning approaches relies on training datasets. These training sets are usually a large database of images that have been manually segmented, which the algorithm is able to “learn” from. Some evidence suggests these algorithms may suffer from overfitting because of the limited dataset that they are trained on, thus leading to poor generalizability of results (17). For example, one study reported significantly lower DSC values when a model trained exclusively on adult CT scans was applied to pediatric CT data, demonstrating degraded segmentation accuracy due to differences in anatomy and image characteristics between training and deployment populations (18).

Objective

The aims of this study were to evaluate the accuracy of a commercially available auto-segmentation algorithm when assessing different joints across various cohorts. We hypothesized that: (I) there would be no difference in auto-segmentation accuracy among different bones within in a joint and (II) auto-segmentation would be more accurate for younger healthy subjects without any known musculoskeletal pathology compared to older subjects or those with known musculoskeletal pathology.


Methods

Participants

This is a secondary analysis of data collected over several research studies focused on the effects of age, injury, and surgical intervention on in vivo joint kinematics.

Data collection

Knee CT scans

Participants were grouped based on age, health status, and diseases or interventions that affect bony knee pathology. Younger healthy knees were knees of adults who had no history of traumatic knee injuries (19,20), patients diagnosed with patellofemoral pain syndrome (PFPS), and the contralateral knees of younger adults who had undergone ACL surgery, with all individuals between the ages of 18 and 44 years (21-23). Younger pathologic knees were from adults who had undergone ACL surgery for grade III injury, with all individuals between the ages of 18 and 41 years (22). The older healthy knees included both older adult subjects’ knees with no pathology and the contralateral knees of older adults undergoing TKA, with all individuals between the ages of 47 and 78 years (24). The older pathologic knees came from older adults undergoing TKA, with all individuals between the ages of 47 and 76 years (24) (Table 1).

Table 1

Summary of demographic and clinical characteristics of subjects (knee scans)

Characteristic Younger healthy knees Younger pathologic knees Older healthy knees Older pathologic knees
Gender, n
Female 35 11 9 6
Male 36 11 17 11
Average age ± SD (years) 26.3±7.0 20.4±3.3 66.0±8.7 64.3±8.3
BMI (kg/m2) 23.6±2.8 24.1±2.2 28.5±4.4 29.1±4.2
Number of knees 71 22 26 17
Average CT resolution (mm) 0.47×0.47×0.47 0.57×0.57×0.57 0.42×0.42×0.42 0.41×0.41×0.41

Data of BMI are presented as mean ± SD. BMI, body mass index; CT, computed tomography; SD, standard deviation.

Shoulder CT scans

Participants were grouped based on age, health status, and history of shoulder pathology. The younger healthy shoulder group comprised adults with no history of shoulder pathology, with all individuals between the ages of 18 and 46 years (25). The older pathological shoulder group included individuals with irreparable rotator cuff tears (defined as tears larger than 5 cm and tears involving more than two tendons) (26) or a history of proximal humerus fractures (PHF). Participants in the PHF cohort had clinically and radiographically healed fractures but demonstrated residual deformity, defined as at least 15 of varus or apex-anterior angulation or displacement of the greater tuberosity of 0.5 cm or more (27). All individuals in the older pathologic group were between the ages of 51 and 79 years. The older healthy group contained contralateral shoulders of people with PHF, with all individuals between the ages of 51 and 79 years (27) (Table 2).

Table 2

Summary of demographic and clinical characteristics of subjects (shoulder scans)

Characteristic Younger
healthy shoulders
Older
pathologic shoulders
Older
healthy shoulders
Gender, n
Female 15 9 7
Male 15 10 3
Average age ± SD (years) 25±7.21 64.55±1.55 66.1±10.08
BMI (kg/m2) 25±4.25 27.84±0.16 27.68±6.00
Number of scans 30 19 10
Average CT resolution (mm) 0.46×0.46×0.46 0.47×0.47×0.47 0.42×0.42×0.42

Data of BMI are presented as mean ± SD. BMI, body mass index; CT, computed tomography; SD, standard deviation.

Data processing

All scans were resliced to isometric voxels before segmentation. All bones were manually segmented using Mimics 24 or 26 software (Materialise NV, Belgium) by technicians who completed training in CT segmentation. Manual segmentation was completed by 4 technicians, and the segmented bones were then visually checked for errors by 2 engineers with 7+ years of experience segmenting medical images. Manual segmentation was completed using a combination of automated thresholding and region growing followed by manual slice-by-slice voxel segmentation. Accuracy of segmented bone tissue was confirmed by registering digitally reconstructed radiographs created from the segmented bone tissue to a series of biplane radiographs using a validated registration process (28,29). The registration process will not reliably converge on a solution if the bones are improperly segmented.

Auto-segmentation of all joints was performed using Simpleware software (version 2024.06; Synopsys, Inc., Sunnyvale View, CA, USA) on a 32 GB RAM, Intel Core i9-7900X CPU @ 3.30 GHz, and NVIDIA GeForce GTX 1080 GPU desktop. The DSC of the entire mask was calculated per slice and averaged for a general DSC per CT scan using a custom MATLAB script (Figure 1B,1C) (8). For the 1-pixel DSC, the 1-pixel edge of the mask per slice was first extracted. Then the same algorithm as the full mask DSC was used to calculate the 1-pixel DSC. The 3D surface models were generated (30) and then a custom MATLAB script computed the mean and standard deviation of the surface deviations (vertex to vertex) (Figure 2). Prior to calculating the accuracy metrics, the MATLAB script cut off the manually segmented bone models at the same height as the automatedly segmented bone models.

Figure 2 Manual segmentation (red) overlayed on auto-segmentation (green). With color map showing surface deviations and histogram showing count of the deviations.

Data analysis

DSC, 1-pixel DSC, and surface deviations were compared using a repeated-measures general linear model with the aligned rank transform to account for unequal variance and non-normality and to evaluate the effects of bone, health status, and age (27). Model assumptions were assessed before inferential testing. Equal variance was evaluated using Levene’s test, and normality was assessed using Shapiro-Wilk and Kolmogorov-Smirnov tests, as well as visual inspection of Q-Q plots. Significant main effects and interactions were further evaluated using aligned rank transform contrasts (ART-C) with Tukey adjustment for pairwise comparisons. Statistical significance was defined as P<0.05. Descriptive statistics, assumption testing, and repeated-measures general linear model procedures were performed using SPSS version 30 (IBM Corp., Armonk, NY, USA), while ART-C post hoc analyses were performed in R version 4.6.0 (R Foundation for Statistical Computing, Vienna, Austria) (31). Effect sizes (ηp2) were calculated and defined as 0.01 (small), 0.06 (medium), and 0.14 (large) (28) to help interpret any observed differences. Additionally, when both knees or both shoulders were healthy in the same participant, only the right-sided scan was retained to preserve independence of observations and avoid within-subject correlation bias. Auto-segmentation time was measured from the time when the “Apply” button was hit after selecting the region of interest to completion of the segmentation process while manual segmentation time was measured only during the segmentation process itself and did not include the registration process described in section Data processing. ChatGPT-4 was used to clarify and condense the results section based on data obtained from R and SPSS. One of the authors (A.O.) validated the data and takes responsibility for accuracy. ChatGPT-4 was also used to write the highlight box based on the final entire manuscript text. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the institutional ethics board of University of Pittsburgh (No. STUDY19060375, STUDY23060157, STUDY21070184, STUDY19080016, STUDY19050059, PRO15070281, STUDY20070400, STUDY23040144, PRO16070246, PRO16050124) and individual consent was obtained from all individual participants.


Results

Knees

A total of 136 knees were segmented (Table 1). Automated segmentation required an average of 50 seconds per unilateral femur and tibia, while manual segmentation required approximately 30 minutes.

DSC

DSC scores in the knee were generally high across groups and bones, although significant differences were observed by group, bone, and the group-by-bone interaction. Overall, the tibia demonstrated significantly higher DSC than the femur (0.967±0.045 vs. 0.933±0.129; P<0.001; Figure 3A), with a large effect size (ηp2=0.343). A significant group effect was also observed (P<0.001; ηp2=0.165; Figure 3B), with DSC values of 0.974±0.026 for older healthy knees, 0.980±0.012 for older pathologic knees, 0.933±0.126 for younger healthy knees, and 0.968±0.016 for younger pathologic knees. ART-C post hoc analysis demonstrated that younger healthy knees had significantly lower DSC than older healthy knees (P<0.001) and older pathologic knees (P<0.001), while younger pathologic knees had significantly lower DSC than older pathologic knees (P=0.044). A significant group-by-bone interaction was also identified (P=0.02; ηp2=0.067; Figure 3C), driven primarily by lower femoral DSC compared with tibial DSC in the younger healthy group (0.905±0.164 vs. 0.960±0.058; P=0.02). In contrast, femoral and tibial DSC values did not differ significantly in the older healthy group (0.971±0.031 vs. 0.977±0.020; P>0.99), older pathologic group (0.977±0.016 vs. 0.983±0.005; P=0.94), or younger pathologic group (0.962±0.019 vs. 0.974±0.009; P=0.84).

Figure 3 The average and standard deviation of (A) the knee bone effect on the DSC, (B) knee group effect on the DSC, (C) Knee bone and group interaction on the DSC, (D) knee bone effect on the 1-pixel DSC, (E) knee group effect on the 1-pixel DSC, (F) knee bone and group interaction on the 1-pixel DSC, (G) knee bone effect on the surface deviation, (H) knee group effect on the surface deviation, (I) knee bone and group interaction on the surface deviation. *, indicates significant differences with P<0.05. DSC, Dice similarity coefficient.

1-pixel DSC

Knee 1-pixel DSC values were poor across groups and bones. Overall, tibial 1-pixel DSC was significantly higher than femoral 1-pixel DSC (0.453±0.157 vs. 0.411±0.176; P=0.008; Figure 3D), although the effect size was small (ηp2=0.046). No statistically significant group effect was observed (P=0.07; ηp2=0.046; Figure 3E), with 1-pixel DSC values of 0.427±0.153 for older healthy knees, 0.376±0.134 for older pathologic knees, 0.441±0.175 for younger healthy knees, and 0.444±0.173 for younger pathologic knees. No significant group-by-bone interaction was observed (P=0.32; ηp2=0.023; Figure 3F), with femoral and tibial values of 0.393±0.134 and 0.462±0.166 in the older healthy group, 0.346±0.147 and 0.407±0.116 in the older pathologic group, 0.426±0.183 and 0.455±0.166 in the younger healthy group, and 0.421±0.204 and 0.466±0.136 in the younger pathologic group.

Surface deviation

Knee surface deviation differed significantly across groups, but not by bone or the group-by-bone interaction. Overall, femoral and tibial surface deviation values were not significantly different (0.336±0.777 vs. 0.283±0.752; P=0.25; Figure 3G), with a negligible effect size (ηp2=0.008). A significant group effect was observed (P<0.001; ηp2=0.142; Figure 3H), with surface deviation values of 0.499±0.418 for older healthy knees, 0.845±1.429 for older pathologic knees, 0.165±0.654 for younger healthy knees, and 0.129±0.654 for younger pathologic knees. ART-C post hoc analysis demonstrated significantly greater surface deviation in older healthy knees compared with younger healthy knees (P<0.001) and younger pathologic knees (P=0.02). Older pathologic knees also demonstrated significantly greater surface deviation than younger healthy knees (P=0.01), while the difference between older pathologic and younger pathologic knees did not reach statistical significance (P=0.09). No significant differences were observed between older healthy and older pathologic knees (P>0.99) or between younger healthy and younger pathologic knees (P>0.99). The group-by-bone interaction was not significant (P=0.08; ηp2=0.042; Figure 3I), with femoral and tibial surface deviation values of 0.608±0.397 and 0.389±0.415 in the older healthy group, 0.904±1.371 and 0.786±1.524 in the older pathologic group, 0.171±0.656 and 0.158±0.655 in the younger healthy group, and 0.062±0.790 and 0.195±0.494 in the younger pathologic group.

Shoulder

A total of 59 shoulders were segmented (Table 2). Automated segmentation averaged 72 seconds, and manual segmentation of the scapula and humerus required approximately 60 minutes.

DSC

Shoulder DSC values were generally high across groups and bones. Overall, humerus demonstrated significantly higher DSC than scapula (0.979±0.029 vs. 0.957±0.022; P<0.001; Figure 4A), with a large effect size (ηp2=0.703). A significant group effect was also observed (P=0.005; ηp2=0.173; Figure 4B), with DSC values of 0.973±0.017 for older healthy shoulders, 0.957±0.040 for older pathologic shoulders, and 0.974±0.018 for younger healthy shoulders. ART-C post hoc analysis demonstrated that older pathologic shoulders had significantly lower DSC than younger healthy shoulders (P=0.003), whereas differences between older healthy and older pathologic shoulders (P=0.29) and older healthy and younger healthy shoulders (P=0.51) were not statistically significant. The group-by-bone interaction was not significant (P=0.10; ηp2=0.079; Figure 4C), with DSC values of 0.986±0.005 for older healthy humerus, 0.959±0.013 for older healthy scapula, 0.966±0.048 for older pathologic humerus, 0.948±0.028 for older pathologic scapula, 0.985±0.007 for younger healthy humerus, and 0.963±0.018 for younger healthy scapula.

Figure 4 The average and standard deviation of (A) the shoulder bone effect on the DSC, (B) shoulder group effect on the DSC, (C) shoulder bone and group interaction on the DSC, (D) shoulder bone effect on the 1-pixel DSC, (E) shoulder group effect on the 1-pixel DSC, (F) shoulder bone and group interaction on the 1-pixel DSC, (G) shoulder bone effect on the surface deviation, (H) shoulder group effect on the surface deviation, (I) shoulder bone and group interaction on the surface deviation. *, indicates significant differences with P<0.05. DSC, Dice similarity coefficient.

1-pixel DSC

Shoulder 1-pixel DSC values were moderate across groups and bones. Overall, scapula demonstrated slightly higher 1-pixel DSC than humerus, although this difference was not statistically significant (0.638±0.168 vs. 0.616±0.173; P=0.37; Figure 4D), with a small effect size (ηp2=0.014). No statistically significant group effect was observed (P=0.15; ηp2=0.067; Figure 4E), with 1-pixel DSC values of 0.654±0.109 for older healthy shoulders, 0.575±0.183 for older pathologic shoulders, and 0.650±0.174 for younger healthy shoulders. The group-by-bone interaction was also not significant (P=0.73; ηp2=0.011; Figure 4F), with humeral and scapula values of 0.647±0.129 and 0.662±0.092 in the older healthy group, 0.552±0.189 and 0.598±0.178 in the older pathologic group, and 0.645±0.170 and 0.656±0.181 in the younger healthy group.

Surface deviation

Shoulder surface deviation did not differ significantly by group, bone, or the group-by-bone interaction. Overall, humerus and scapula surface deviation values were not significantly different (0.060±0.388 vs. 0.023±0.393; P=0.40; Figure 4G), with a small effect size (ηp2=0.013). No statistically significant group effect was observed (P=0.56; ηp2=0.021; Figure 4H), with surface deviation values of 0.081±0.357 for older healthy shoulders, −0.054±0.434 for older pathologic shoulders, and 0.089±0.365 for younger healthy shoulders. The group-by-bone interaction was also not significant (P=0.57; ηp2=0.020; Figure 4I), with humeral and scapula surface deviation values of 0.146±0.360 and 0.016±0.361 in the older healthy group, −0.006±0.469 and −0.102±0.402 in the older pathologic group, and 0.073±0.345 and 0.105±0.388 in the younger healthy group.


Discussion

Overall performance (DSC and surface deviation)

This study evaluated the accuracy of a commercially available auto-segmentation algorithm across multiple patient populations and anatomical regions using three complementary validation metrics based on the gold standard of manual segmentation. For the knee, mean DSC values ranged from 0.905 to 0.983 for the femur and tibia. These values are slightly lower than those previously reported in the literature for automated tibial and femoral segmentation (DSC >0.96) (1). Mean surface deviation for knee segmentation (0.062–0.904 mm) was within the 1-mm range reported as potentially useful for preoperative planning (14).

For the shoulder, mean DSC values ranged from 0.957 to 0.979 for the scapula and humerus, consistent with literature values (>0.94) considered to represent good segmentation performance (9,32). In addition, the range of surface deviations observed in the present study fall within the range of previously reported acceptable accuracy thresholders of 1–2.5 mm and are therefore considered sufficiently accurate for clinical applications (9).

Overall, the automated segmentation approach demonstrated acceptable accuracy across all bones while substantially reducing segmentation time, achieving approximately 36-fold, and 50-fold time savings for the knee and shoulder respectively, compared with manual segmentation.

Performance across bone types

We hypothesized that segmentation accuracy would not differ among bones within a joint; however, this hypothesis was not supported. In the knee, the tibia achieved significantly higher DSC than the femur, with a large effect size. This finding contrasts with reports in the literature, where the femur generally shows better performance than the tibia (1). In the shoulder, scapula segmentations showed lower DSC than humeral segmentations, again with a large effect size. This likely reflects the scapula’s more complex geometry, including the thin subscapula fossa and lateral borders, which are often only one to two voxels thick (33). Similar findings have been reported in literature with scapula performing worse (32). Notably, surface deviation did not differ significantly between bones in either of the two joints, indicating the small (yet statistically significant) differences in DSC do not compromise the accuracy of the three-dimensional models created from the segmentations. Together, these results demonstrate that segmentation accuracy is inherently bone-dependent and that anatomical geometry plays a critical role in algorithm performance, underscoring the importance of bone-specific validation rather than assuming uniform accuracy across all bones within a joint.

Performance across patient populations

We further hypothesized that auto-segmentation accuracy would be best in younger healthy subjects; however, this hypothesis was not consistently supported across joints. In the knee, the opposite trend was observed: knees of older healthy and older pathologic cohorts demonstrated significantly higher DSC than knees in younger healthy subjects, while segmentation of knees from younger pathologic patients was worse than in knees of older pathologic patients. This population effect was substantial, with a large effect size, indicating a meaningful influence of age and disease status on segmentation performance in the knee. We can speculate that this finding may be due to the training dataset for the evaluated automated knee segmentation primarily comprising of older patients. Similar findings have been reported in the literature, where deep learning-based segmentation of knee articular cartilage from 3D ultrasound images performed better in osteoarthritic knees than in healthy ones, emphasizing the role of the training dataset (34). In contrast, surface deviation increased from younger healthy and younger pathologic to older healthy, with a large effect size, reflecting greater boundary-level variability with advancing age and pathology despite high volumetric overlap. This result is supported by the fact that changes in bone mineral density, narrowing of joint spaces and presence of irregular osteophytes may disrupt an automated algorithms ability to accurately segment the boundaries of bones (35).

In the shoulder, the hypothesis was partially supported, as shoulders in the older pathologic group exhibited significantly lower DSC than younger healthy subjects, with a large effect size, although no corresponding differences in surface deviation were detected. This difference may be because our older pathologic group included individuals with conservatively managed PHF containing a bony callus which may affect the bone shape and density in certain regions. Collectively, these findings indicate that the effect of age and pathology on segmentation accuracy is joint-specific, exerting the greater influence in the knee than in the shoulder.

1-pixel DSC

Across the shoulder and knee datasets, 1-pixel DSC values did not differ significantly by age group or pathologic status, and no bone-by-group interactions were observed. However, knee 1-pixel DSC was significantly higher for the tibia than the femur. These results indicate that the auto-segmentation algorithm achieves consistently low to moderate boundary-level performance across anatomical regions and clinical populations, including in the presence of degenerative or traumatic pathology. Notably, the absence of significant differences between pathologic and healthy joints suggests that abnormal surface features, such as osteophytes or fracture-related irregularities, did not disproportionately affect edge localization.

These findings highlight important limitations of conventional DSC when used in isolation. Because DSC is dominated by agreement within the interior volume of a structure, it is relatively insensitive to small boundary errors, particularly in large bones where clinically meaningful surface inaccuracies may be masked (36). This limitation is especially consequential for thin tissues such as articular cartilage, which may be only several pixels thick on clinical imaging (37); in this setting, a 1 to 2 pixel surface error can represent a substantial proportion of the tissue thickness despite a high overall DSC. Prior cartilage magnetic resonance imaging (MRI) auto-segmentation studies similarly show that high volumetric overlap does not necessarily reflect accurate boundary delineation or thickness estimation (38). In contrast, 1-pixel DSC directly evaluates boundary agreement and therefore provides complementary information more relevant to applications requiring precise cortical geometry, including joint contact modeling, implant design, and 2D–3D registration of bone models to radiographs. Collectively, these results support using 1-pixel DSC as an adjunct to conventional metrics when validating segmentation algorithms for high-precision orthopaedic and biomechanical applications.

Strengths and limitations

A major strength of this study is the availability of known bone morphology for both healthy and pathologic joints, which enabled objective evaluation of segmentation performance across clinically relevant conditions. In addition, our cohort encompassed diverse clinical populations spanning multiple age groups and pathologic states, allowing assessment of the auto-segmentation algorithm across a wide range of individuals and stages of joint degeneration.

Importantly, this study benefited from a high-quality reference standard, which is critical in external validation of segmentation algorithms. The manually segmented bone models used for comparison were not only generated in large numbers, but were also validated through successful registration to biplane radiographs. Because the registration algorithm does not reliably converge when bone geometry is inaccurate, consistent and accurate registration provides strong confirmation of the fidelity of the manual segmentations. This strengthens confidence that the observed performance of the auto-segmentation method reflects true algorithmic accuracy rather than uncertainty in the ground truth.

Several limitations should be acknowledged. First, the sample size was relatively small for some groups, which may limit generalizability. Second, our findings are restricted to the segmentation of osseous tissue from CT imaging and should not be extrapolated to soft-tissue segmentation tasks (e.g., cartilage or intervertebral discs) using MRI. Next, only one commercially available software was evaluated, and the results are likely highly dependent upon the data used to train the algorithm, so these results should not be extrapolated other auto-segmentation algorithms. Simpleware software uses a proprietary machine-learning auto-segmentation workflow. Therefore, the present study should be interpreted as an external evaluation of the commercially available implementation rather than as validation of a fully specified open model. No intra-class correlation coefficient was calculated to evaluate agreement between technicians. Technicians were not blinded to patient grouping. Finally, the results are specific to the bones and joints evaluated in this study and may not extend to more anatomically complex structures, such as vertebrae.


Conclusions

The commercially available auto-segmentation algorithm evaluated in this study demonstrated generally high accuracy for segmenting knee and shoulder bones from CT scans while providing substantial time savings compared with manual segmentation. The performance of the algorithm was influenced by age and pathology in a joint- and metric-specific manner: in the knee, older subjects showed higher DSC but greater surface deviation than younger subjects, regardless of pathology, with no age- or pathology-related differences in 1-pixel DSC; in the shoulder, older pathologic subjects showed lower DSC than younger healthy subjects, while surface deviation and 1-pixel DSC did not differ by age or pathology. Whole-bone volumetric agreement was consistently strong across bones, joints, and patient populations, although boundary-level accuracy remained comparatively limited. Overall, these findings suggest that the evaluated auto-segmentation algorithm is a reliable and efficient alternative to manual segmentation for applications that do not require precise identification of small cortical defects or fine surface abnormalities. However, for high-precision tasks that depend on accurate cortical edge definition, additional validation or refinement of automated methods may be necessary.

The authors would like to acknowledge Sabreen Megherhi, Zhaoyi Fang, Gillian Kane, Nathan Hyre for helping with data processing. Additionally, the data included in this manuscript was presented as a poster at American Society of Biomechanics in 2024 and Orthopedic Research Society in 2025. ChatGPT-4 was used to clarify and condense the results section based on data obtained from R and SPSS. One of the authors (A.O.) validated the data and takes responsibility for accuracy. ChatGPT-4 was also used to write the highlight box.


Footnote

Data Sharing Statement: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0059/dss

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0059/prf

Funding: This work was supported in part by NIH/NIAMS and Research Support from Smith & Nephew (Grants Nos. R44 AR064620 and R01 AR080425).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2026-0059/coif). W.A. received the grant from NIH and Smith & Nephew paid to university. The other authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the institutional ethics board of University of Pittsburgh (No. STUDY19060375, STUDY23060157, STUDY21070184, STUDY19080016, STUDY19050059, PRO15070281, STUDY20070400, STUDY23040144, PRO16070246, PRO16050124) and individual consent was obtained from all individual participants.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Humayun A, Rehman M, Liu B. A method framework of semi-automatic knee bone segmentation and reconstruction from computed tomography (CT) images. Quant Imaging Med Surg 2024;14:7151-75. [Crossref] [PubMed]
  2. Chang Y, Yuan Y, Guo C, et al. Accurate Pelvis and Femur Segmentation in Hip CT With a Novel Patch-Based Refinement. IEEE J Biomed Health Inform 2019;23:1192-204. [Crossref] [PubMed]
  3. Dauwe J, Mys K, Putzeys G, et al. Advanced CT visualization improves the accuracy of orthopaedic trauma surgeons and residents in classifying proximal humeral fractures: a feasibility study. Eur J Trauma Emerg Surg 2022;48:4523-9. [Crossref] [PubMed]
  4. Saillard E, Gardegaront M, Levillain A, et al. Finite element models with automatic computed tomography bone segmentation for failure load computation. Sci Rep 2024;14:16576. [Crossref] [PubMed]
  5. Song P, Fan Z, Zhi X, et al. Study on the accuracy of automatic segmentation of knee CT images based on deep learning. Zhongguo Xiu Fu Chong Jian Wai Ke Za Zhi 2022;36:534-9. [Crossref] [PubMed]
  6. Kuiper RJA, Sakkers RJB, van Stralen M, et al. Efficient cascaded V-net optimization for lower extremity CT segmentation validated using bone morphology assessment. J Orthop Res 2022;40:2894-907. [Crossref] [PubMed]
  7. Almeida DF, Ruben RB, Folgado J, et al. Fully automatic segmentation of femurs with medullary canal definition in high and in low resolution CT scans. Med Eng Phys 2016;38:1474-80. [Crossref] [PubMed]
  8. Falcinelli C, Cheong VS, Ellingsen LM, et al. Segmentation methods for quantifying X-ray Computed Tomography based biomarkers to assess hip fracture risk: a systematic literature review. Front Bioeng Biotechnol 2024;12:1446829. [Crossref] [PubMed]
  9. Satir OB, Eghbali P, Becce F, et al. Automatic quantification of scapular and glenoid morphology from CT scans using deep learning. Eur J Radiol 2024;177:111588. [Crossref] [PubMed]
  10. Noguchi S, Nishio M, Yakami M, et al. Bone segmentation on whole-body CT using convolutional neural network with novel data augmentation techniques. Comput Biol Med 2020;121:103767. [Crossref] [PubMed]
  11. Fu Y, Liu S, Li HH, et al. Automatic and hierarchical segmentation of the human skeleton in CT images. Phys Med Biol 2017;62:2812-33. [Crossref] [PubMed]
  12. Whittier DE, Boyd SK, Burghardt AJ, et al. Guidelines for the assessment of bone density and microarchitecture in vivo using high-resolution peripheral quantitative computed tomography. Osteoporos Int 2020;31:1607-27. [Crossref] [PubMed]
  13. Mahfouz MR, Hoff WA, Komistek RD, et al. Effect of segmentation errors on 3D-to-2D registration of implant models in X-ray images. J Biomech 2005;38:229-39. [Crossref] [PubMed]
  14. Bori E, Pancani S, Vigliotta S, et al. Validation and accuracy evaluation of automatic segmentation for knee joint pre-planning. Knee 2021;33:275-81. [Crossref] [PubMed]
  15. Liang P, Li Y, Feng P, et al. Advancing bone tumor detection in older adults: the impact of AI-enhanced medical imaging. Front Med (Lausanne) 2025;12:1697975. [Crossref] [PubMed]
  16. Engelkes K. Accuracy of bone segmentation and surface generation strategies analyzed by using synthetic CT volumes. J Anat 2021;238:1456-71. [Crossref] [PubMed]
  17. Lee SB, Hong Y, Cho YJ, et al. Deep Learning-Based Computed Tomography Image Standardization to Improve Generalizability of Deep Learning-Based Hepatic Segmentation. Korean J Radiol 2023;24:294-304. [Crossref] [PubMed]
  18. Kumar K, Yeo AU, McIntosh L, et al. Deep Learning Auto-Segmentation Network for Pediatric Computed Tomography Data Sets: Can We Extrapolate From Adults? Int J Radiat Oncol Biol Phys 2024;119:1297-306. [Crossref] [PubMed]
  19. Xu C, Aloi N, Gale T, et al. Symmetry in knee arthrokinematics in healthy collegiate athletes during fast running and drop jump revealed through dynamic biplane radiography. Osteoarthritis Cartilage 2023;31:1501-14. [Crossref] [PubMed]
  20. Gale T, Anderst W. Knee Kinematics of Healthy Adults Measured Using Biplane Radiography. J Biomech Eng 2020;142:101004. [Crossref] [PubMed]
  21. Nishida K, Gale T, Chiba D, et al. The effect of lateral extra-articular tenodesis on in vivo cartilage contact in combined anterior cruciate ligament reconstruction. Knee Surg Sports Traumatol Arthrosc 2022;30:61-70. [Crossref] [PubMed]
  22. Gibbs CM, Hughes JD, Popchak AJ, et al. Anterior cruciate ligament reconstruction with lateral extraarticular tenodesis better restores native knee kinematics in combined ACL and meniscal injury. Knee Surg Sports Traumatol Arthrosc 2022;30:131-8. [Crossref] [PubMed]
  23. Gibbs CM, Hughes JD, Popchak AJ, et al. Preoperative quantitative pivot shift does not correlate with in vivo kinematics following ACL reconstruction with or without lateral extraarticular tenodesis. Knee Surg Sports Traumatol Arthrosc 2023;31:2802-9. [Crossref] [PubMed]
  24. Copp EH, Gale TH, Byrapogu VKC, et al. Unicompartmental knee arthroplasty approximates healthy knee kinematics more closely than total knee arthroplasty. J Orthop Res 2024;42:2514-24. [Crossref] [PubMed]
  25. Brown CL, LeVasseur CM, Scott D, et al. Best-Fit Circle Missing Area Method Shows Good Accuracy and Interrater Reliability When Assessing Glenoid Bone Loss. Am J Sports Med 2025;53:2060-5. [Crossref] [PubMed]
  26. LeVasseur CM, Kane G, Hughes JD, et al. In Vivo Graft Elongation After Arthroscopic Dermal Superior Capsular Reconstruction. Am J Sports Med 2023;51:2671-8. [Crossref] [PubMed]
  27. Kuhn MZ, Sudar A, LeVasseur C, et al., editors. Varus Malunion of Proximal Humerus Fracture Alters Glenohumeral Joint Kinematics. ORS 2025 Annual Meeting; 2025; Phoenix, AZ.
  28. Anderst W, Zauel R, Bishop J, et al. Validation of three-dimensional model-based tibio-femoral tracking during running. Med Eng Phys 2009;31:10-6. [Crossref] [PubMed]
  29. Bey MJ, Zauel R, Brock SK, et al. Validation of a new model-based tracking technique for measuring three-dimensional, in vivo glenohumeral joint kinematics. J Biomech Eng 2006;128:604-9. [Crossref] [PubMed]
  30. Treece GM, Prager RW, Gee AH. Regularised marching tetrahedra: improved iso-surface extraction. Computers & Graphics 1999;23:583-98.
  31. Elkin LA, Kay M, Higgins JJ, et al. An Aligned Rank Transform Procedure for Multifactor Contrast Tests. The 34th Annual ACM Symposium on User Interface Software and Technology; Virtual Event, USA: Association for Computing Machinery; 2021. p. 754-68.
  32. Sezer A, Sezer A. Mask Region-Based Convolutional Neural Network segmentation of the humerus and scapula from proton density-weighted axial shoulder magnetic resonance images. Jt Dis Relat Surg 2023;34:583-9. [Crossref] [PubMed]
  33. Yang Z, Fripp J, Chandra SS, et al. Automatic bone segmentation and bone-cartilage interface extraction for the shoulder joint from magnetic resonance images. Phys Med Biol 2015;60:1441-59. [Crossref] [PubMed]
  34. du Toit C, Orlando N, Papernick S, et al. Automatic femoral articular cartilage segmentation using deep learning in three-dimensional ultrasound images of the knee. Osteoarthr Cartil Open 2022;4:100290. [Crossref] [PubMed]
  35. Rossi M, Marsilio L, Mainardi L, et al. CEL-Unet: Distance Weighted Maps and Multi-Scale Pyramidal Edge Extraction for Accurate Osteoarthritic Bone Segmentation in CT Scans. Front Signal Process 2022;2:857313.
  36. Mohammadi M, Mollazade K, Behroozi-Khazaei N. Under- and over-segmentation: New metrics for image segmentation accuracy measurement. Array 2025;28:100624.
  37. Joseph GB, Baum T, Alizai H, et al. Baseline mean and heterogeneity of MR cartilage T2 are associated with morphologic degeneration of cartilage, meniscus, and bone marrow over 3 years--data from the Osteoarthritis Initiative. Osteoarthritis Cartilage 2012;20:727-35. [Crossref] [PubMed]
  38. Desai A, Caliva F, Iriondo C, et al. A multi-institute automated segmentation evaluation on a standard dataset: Findings from the international workshop on osteoarthritis imaging segmentation challenge. Osteoarthritis and Cartilage 2020;28:S304-5.
doi: 10.21037/jmai-2026-0059
Cite this article as: Olawin A, Gale T, Gray EC, LeVasseur C, Ramraj R, Krakora E, Lin A, Urish K, Moloney G, Hogan M, Anderst W. Comparison of machine-learning-based auto-segmentation to manual segmentation of knee and shoulder CT scans. J Med Artif Intell 2026;09:70.

Download Citation