Model development and validation for classifying hypoxia in military aircrew using ECG and skin temperature
Original Article

Model development and validation for classifying hypoxia in military aircrew using ECG and skin temperature

Maarten P. D. Schadd1 ORCID logo, Jan Ubbo van Baardewijk2 ORCID logo, Mattia Tarchini Bojczuk3 ORCID logo, Alessandro Aliberti3,4 ORCID logo, Fred L. Vuijk2, Lotte Linssen2 ORCID logo, Kaj Gijsbertse2 ORCID logo, Mark M. J. Houben2 ORCID logo, Mario Arrigoni-Neri5,6 ORCID logo, Boris R. M. Kingma2 ORCID logo, Eugene P. van Someren7 ORCID logo

1Intelligence and Decision Support, Netherlands Organisation for Applied Scientific Research (TNO), The Hague, The Netherlands; 2Human Performance, Netherlands Organisation for Applied Scientific Research (TNO), Soesterberg, The Netherlands; 3AlphaWaves S.r.l., Torino, Italy; 4Inter-university Dept. of Regional and Urban Studies and Planning, Politecnico di Torino, Torino, Italy; 5QM Project, Via G.B. Pioda 14, Lugano, Switzerland; 6Politecnico di Milano, Milan, Italy; 7Risk Analysis for Prevention, Innovation & Development, TNO, Utrecht, The Netherlands

Contributions: (I) Conception and design: MPD Schadd, JU van Baardewijk, EP van Someren; (II) Administrative support: MMJ Houben, BRM Kingma; (III) Provision of study materials or patients: FL Vuijk, L Linssen, K Gijsbertse; (IV) Collection and assembly of data: FL Vuijk, L Linssen, K Gijsbertse; (V) Data analysis and interpretation: M Tarchini Bojczuk, A Aliberti, MPD Schadd, JU van Baardewijk, M Arrigoni-Neri; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Maarten P. D. Schadd, PhD. Intelligence and Decision Support, Netherlands Organisation for Applied Scientific Research (TNO), P.O. Box 96864, 2509 JG The Hague, The Netherlands. Email: maarten.schadd@tno.nl.

Background: Hypoxia occurs when blood or tissues are deprived of adequate oxygen, posing a significant risk to military aircrew operating at high altitudes due to reduced atmospheric pressure. The danger lies in its subtle symptoms—such as impaired judgment—often unnoticed until serious consequences arise. While hypoxia can be detected using direct (e.g., pulse oximetry), indirect (e.g., photoplethysmography), or tissue-level (e.g., near-infrared spectroscopy) methods, these are often impractical in flight settings due to motion artifacts, low perfusion, or invasiveness. This study aims to develop and internally validate machine learning models to classify hypoxic conditions in military aircrew members using electrocardiogram (ECG) and skin temperature signals, offering a non-invasive and real-time monitoring approach.

Methods: Data were collected from healthy military aircrew members undergoing standardized hypoxia training in a hypobaric chamber simulating high-altitude conditions. ECG, skin temperature, and respiration signals were recorded using wearable sensors. Hypoxia events were labeled based on oxygen mask removal at altitude. A multi-window feature extraction approach was applied using time windows of 20–120 seconds, enabling trend detection across scales. In total, 87 ECG-derived features were combined with temperature and respiration features. After preprocessing and quality control, a feature selection strategy based on repeated classifier rankings was employed. Five classifiers were trained and evaluated (support vector classification, decision tree, random forest, logistic regression, K-nearest neighbors) using cross validation and accuracy and log loss metrics.

Results: Data from 40 participants were included after preprocessing. Across classifiers, classification accuracy ranged from 0.85 to 0.90, with the support vector classification achieving the highest average accuracy (0.90±0.07) and lowest log loss (0.25±0.15). The most informative features came from longer time windows, particularly the 80th percentile of respiratory rate intervals (HRV_Prc80NN), mean respiration rate, and mean skin temperature. Classifier performance was robust across models, with small differences, suggesting that model architecture is less critical than feature representation.

Conclusions: We demonstrate that ECG and skin temperature signals can reliably detect hypoxic conditions in military aircrew using machine learning. The approach shows strong internal performance and highlights specific physiological features as key indicators. These findings support the feasibility of real-time, non-invasive hypoxia monitoring in flight environments and lay the groundwork for future applications in other high-risk domains such as commercial aviation, spaceflight, and clinical monitoring. Further research will involve evaluating model performance in operational flight conditions and exploring generalization across individuals and sensor systems.

Keywords: Hypoxia; electrocardiogram (ECG); skin temperature; machine learning; feature importance


Received: 09 March 2025; Accepted: 01 July 2025; Published online: 11 September 2025.

doi: 10.21037/jmai-2024-469


Highlight box

Key findings

• For classifying hypoxic conditions of military aircrew members, this study found a high feature importance of electrocardiogram (ECG)-derived heart-rate variability, ECG-derived respiratory rate intervals and skin temperature. Support vector classification performs best with an accuracy of 90% on a balanced dataset.

What is known and what is new?

• Machine learning methods have been used to classify hypoxia based on ECGs before, but primarily in fetal or intensive care unit settings and often using only 3–10 features from the ECG. In contrast, we have leveraged a unique dataset that accurately simulates hypoxia exposure during training of military aircrew personnel. In addition, we employed a comprehensive set of 87 ECG measures and skin temperature measures in six-time windows.

What is the implication, and what should change now?

• We have shown that hypoxia can be predicted in a realistic setting from an ECG and skin temperature sensor, which opens the way to advance this approach towards practical applications to support military aircrew members in the cockpit. In addition, more research is warranted for extending the amount of sensors towards classifying different types of human performance factors.


Introduction

Background

Aircrew operating at high altitudes are at risk of hypoxia, a condition in which insufficient oxygen reaches body tissues (1,2). The reduced atmospheric pressure encountered during flight makes it harder for oxygen to move from the lungs into the blood, leading to physiological and cognitive deterioration (2-7). Despite its dangers, hypoxia is difficult for individuals to detect because it typically causes no pain or discomfort and impairs self-awareness (8). Early effects may include subtle visual impairments such as reduced color vision, contrast sensitivity, and peripheral light awareness (9-12), while prolonged or severe oxygen deprivation can lead to rapid breathing, cyanosis, mental confusion, loss of coordination, and, ultimately, unconsciousness (2,3). These symptoms vary depending on altitude, individual sensitivity, and the rate and duration of exposure (2,13).

To prepare for such conditions, military aircrew undergo mandatory hypoxia awareness training in hypobaric chambers that simulate altitude effects in a controlled environment (14). These sessions also present a valuable opportunity for collecting realistic physiological data to improve detection methods.

Various technologies are used to monitor hypoxia. Pulse oximetry, which measures blood oxygen saturation (SpO2), is widely adopted due to its non-invasive nature (15-17), but its reliability can be compromised by low perfusion, skin pigmentation, motion, and rapid altitude changes (18,19). Other approaches assess oxygen consumption indirectly, such as breath gas analysis and photoplethysmography (PPG). Breath analysis systems (e.g., Oxycon Mobile) capture variables such as oxygen uptake (VO2) and the volume of carbon dioxide output (VCO2) to monitor respiratory function (20-23), though their equipment requirements limit real-world usability. PPG offers a wearable alternative by monitoring blood volume changes via light sensors (24), but peripheral vasoconstriction and motion artifacts can affect signal quality (25,26).

Tissue-level oxygenation monitoring, using methods such as near-infrared spectroscopy (NIRS), provides insight into local oxygen availability and metabolic activity (27-29). However, sensor placement constraints, combined with environmental stressors such as high G-forces, can limit their applicability for pilots in flight.

Given these challenges, electrocardiogram (ECG)-based detection of hypoxia, particularly via heart rate variability (HRV) analysis, has gained interest as a promising, unobtrusive alternative (30,31). Prior studies have demonstrated links between HRV and hypoxia in both general and aviation-specific populations (32-34), supported by the growing accessibility of HRV measurement (35,36). While machine learning models have been explored for ECG-based hypoxia detection, they have mostly focused on fetal monitoring (37-39) or intensive care unit (ICU) settings (40), highlighting a gap in research focused on real-world, in-flight conditions.

Rationale and knowledge gap

Despite growing interest in physiological monitoring for hypoxia detection, current research presents several critical limitations. Many prior studies have focused on fetal hypoxia (37-39), which limits their applicability to adult populations, particularly healthy, trained aircrew. Other studies, such as those by Niaussat et al. (40), use clinical data from intensive care patients, which do not reflect the high-performance context of aviation operations. Furthermore, while relationships between hypoxia and HRV have been explored (32-34), existing work typically investigates a narrow selection of features, often without assessing how these features behave across different time windows or under realistic conditions relevant to flight.

Another gap lies in the nature of the datasets used. For example, the Harespod dataset (41) consists of physiological recordings from college students who underwent controlled hypoxia exposure via gradual decompression in a hypobaric chamber. While valuable, this dataset does not fully replicate the operational hypoxia conditions encountered by trained aircrew, limiting its applicability to military aviation contexts. Additionally, few studies have collected data under sudden decompression, a more realistic simulation of in-flight hypoxic incidents.

To address these gaps, our work introduces two key advancements. First, we conduct a comprehensive analysis of ECG-derived features, covering a broad range of metrics across multiple temporal windows; an approach that allows for deeper insights into feature effectiveness. Second, we utilize a novel dataset collected during hypobaric chamber training of military aircrew under sudden decompression, representing one of the most ecologically valid datasets available for studying in-flight hypoxia detection.

Objective

The objective of this study is to develop and evaluate a machine-learning-based model for hypoxia classification in military aircrew members using an extensive set of 87 ECG-based features combined with three non-ECG features. We present this article in accordance with the TRIPOD reporting checklist (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2024-469/rc).


Methods

Dataset description

This section describes how we collected hypoxia data of military aircrew members. This dataset forms the basis for our machine-learning experiments.

Hypobaric chamber

A hypobaric chamber, also known as an altitude chamber, is an enclosed space used to simulate high-altitude conditions by reducing the air pressure inside the chamber (42,43). This reduced pressure environment replicates the atmospheric conditions found at various altitudes above sea level during a flight. At higher altitudes less oxygen is available, which leads to impaired cognitive and physical functions, negatively influencing flight performance. During hypoxia training, aircrew members will learn how to recognize the symptoms of hypoxia and how to respond accordingly to prevent in-flight emergency situations. This training helps them recognize and respond to hypoxia.

The training usually involves a group of aircrew members seated within the chamber, allowing them to experience hypoxic conditions together under observation. An exercise controller, usually an experienced flight physiologist or medical professional, oversees the session from outside the chamber. They monitor the aircrew members’ physiological responses to the decompression, communicate with them, and manage the chamber’s conditions.

The session begins at ground-level pressure, allowing aircrew members to acclimate to the environment and prepare for the upcoming decompression. Once ready, the chamber undergoes a sudden reduction in pressure, rapidly simulating the altitude of 25,000 feet. As the oxygen levels drop, aircrew members are encouraged to recognize the physical and cognitive symptoms of hypoxia in themselves and observe the signs in their peers (see “Background”).

Throughout the exercise, the controller prompts aircrew members to complete simple tasks or respond to questions (e.g., completing a simple maze, date recollection), helping them observe how hypoxia affects their performance. After this section, an altitude of 18,000 feet is simulated, wherein the aircrew members observe the effects of hypoxia on night-vision. By the end of the session, aircrew members are provided oxygen masks, and the chamber is gradually returned to normal pressure. This exposure prepares aircrew members to recognize hypoxia symptoms in flight and respond quickly to avoid loss of consciousness or other severe consequences.

Sensors for collecting data

The participants were equipped with a custom designed ‘physiology monitoring sensor’, called the ‘Health Patch’ [2M Engineering, Valkenswaard, The Netherlands (44)]. This sensor was attached to the body of aircrew members as depicted in Figure 1.

Figure 1 Sensor locations of the physiology monitoring system. ECG, electrocardiogram; PPG, photoplethysmography; SpO2, blood oxygen saturation.

This sensor included multiple components strategically placed on the participants to capture distinct physiological signals. The specific sensors and their corresponding data types include:

  • ECG, at a sampling frequency of 250 Hz. The electrodes (3M Red Dot, type 2330) were placed on the participants’ chest (see also Figure 1), arranged to provide accurate and reliable heart rate data without interference of other flight equipment.
  • Respiratory rate (RR) derived from the bioimpedance signal (45). The bioimpedance signal was retrieved through the same electrodes as the ECG. These sensors detect changes in the body’s electrical impedance caused by breathing movements (i.e., lungs filled with oxygen have a higher impedance).
  • PPG data, which measures blood volume changes in the microvascular bed of tissue. This was collected using an optical sensor that was positioned on the forehead. The PPG data was used to derive the SpO2.
  • Skin temperature. Temperature sensors were placed on the forehead.
  • Pressure in the hypobaric chamber (its inverse is referred to as altitude).

The skin temperature is useful as hypoxia may induce changes in the body’s thermoregulatory mechanisms. As the body tries to compensate for reduced oxygen levels, blood flow and skin temperature may change. The respiration rate is a useful feature as one of the immediate body compensatory responses is to increase oxygen intake by increasing the respiration rate.

Next to these automatic data gathering, two auxiliary data categories were collected manually. The first are the timestamps when the aircrew members placed their mask on/off their face. This is relevant, as even when the air pressure is low aircrew members cannot become hypoxic when they are wearing an oxygen mask. The second are timestamps of the beginning and end of the experiment, so that irrelevant data can easily be filtered out.

Data collection protocol

Fifty-one aircrew members undergoing their regular hypoxia training at the Center for Man in Aviation (CMA) in Soesterberg, the Netherlands, were included in this study (see Table 1).

Table 1

Baseline characteristics of participants by sex

Sex Age (years) Weight (kg) Height (cm)
Mean (SD) Range Mean (SD) Range Mean (SD) Range
Female (N=8) 28.0 (2.5) 25–32 68.5 (11.2) 50–82 170.2 (5.7) 160–176
Male (N=43) 30.4 (6.7) 21–54 87.6 (8.7) 66–104 186.0 (6.2) 170–200

SD, standard deviation.

Data collection was conducted during nine training sessions in 2023, specifically on June 5th, August 21st, 28th, 29th and 31st, October 19th and 26th, and November 6th and 7th. Training sessions were performed between 1 and 4 PM, and typically 6–10 aircrew members were trained in the same session. Before starting their hypoxia training in the hypobaric chamber, each participant was fitted with three ECG electrodes, a PPG sensor, and a signal recorder by the experimental leader as described in “Sensors for collecting data”. During the training, the experimental leader monitored participants from the outside, logging when masks were put on and off.

Hypoxia training started with 30 minutes of acclimatization at sea level, which included a gradual altitude ascent to clear ears. After around 30 minutes, the participants experienced a rapid decompression to 25,000 feet, sequentially removing their masks to induce hypoxia, with a maximum exposure of six minutes. The exposure to hypoxia was ended by the participant him-/herself or by the instructor. The procedure was repeated at 18,000 feet in darkness, with masks reapplied after four minutes. The session concluded with a gradual descent.

After carefully removing the sensors, the collected data was downloaded and organized by participant number and training day. The experimental protocol was conducted in accordance with the Declaration of Helsinki and its subsequent amendments and was approved by the Research Ethics Board of TNO (2023-030). All participants provided free and informed consent prior to participating in this study.

Ground truth labels

By assigning labels, we essentially create a “ground truth” that allows the model to learn the relationships between these features and when an aircrew member is hypoxic. For our study, this involves labeling segments of ECG data as either “Hypoxic” or “Nonhypoxic” to represent the subject’s oxygenation state. The labeling strategy has a profound impact on the machine-learning model that tries to replicate the labels.

Measuring SpO2 would normally be used for measuring oxygenation. However, these sensors are not validated for use in low pressure environments such as the hypobaric chamber, and preliminary experiments have shown fluctuating signals. Therefore, a different labeling strategy was chosen. All timestamps at which the altitude was less than 10,000 feet or at which the participant was wearing an oxygen mask were labeled as nonhypoxic. The reasoning is that an aircrew member cannot be hypoxic when not at altitude or when receiving oxygen through the mask. For the remainder of the signal, we use the manually recorded mask on/off events. ECG timestamps are labeled as hypoxic when the mask is off and air pressure is below 50% of its normal. In all other cases the ECG signal is marked as nonhypoxic.

The advantages are that mask-off event markers are non-fluctuating, reliable and subject independent (as they are manually recorded and not dependent on a sensor). Air pressure can also be measured accurately. A potential disadvantage is possible transition-state segments, where the physiological response in the ECG may still be nonhypoxic despite the mask being off in low air pressure. While we find it worthwhile to validate this labeling approach against direct physiological markers of hypoxia, such validation was beyond the scope of the present study.

For the remainder of the article, when referring to hypoxic or nonhypoxic data points, it is implied that for hypoxic data points the mask is off and air pressure is below 50% of its normal value.

Data exploration

To provide an overview of the experimental session and illustrate data quality, we present examples in Figures S1-S3. Each figure displays SpO2, mask-off events, skin temperature, and pressure data from a single participant during the experiment.

The experiment session for each subject includes two mask-off events: a first one at 25,000 feet altitude and a second mask-off event at 18,000 feet altitude. During the first and second mask-off events at high altitude, participants are subjected to reduced oxygen levels for a maximum of 6 minutes, leading to hypoxia. This condition is shown by changes in SpO2 levels. For instance, the SpO2 levels of the subject shown in Figure S1 show a significant decrease during hypoxia exposure at 25,000 feet and a lesser, though notable, decrease at 18,000 feet. For example, two subjects that exhibit good quality SpO2 readings are illustrated in Figures S1,S3. In contrast, one subject that shows highly fluctuating SpO2 signals is illustrated in Figure S3. In general, low pressure combined with mask-off conditions corresponds to the most pronounced decreases in SpO2, consistent with exposure to hypoxic environments. However, even when pressure is high and the mask is on, SpO2 signals can be noisy or unstable, suggesting that signal quality may be influenced by factors other than oxygen availability alone. Additionally, no clear or consistent changes in skin temperature are observed immediately following sudden decompression events.

In the end, 40 out of 51 participants were included for further data processing (7 female, 33 male participants). Eleven participants were excluded due to incomplete data or subjects dropping out early in the training session (see also Figure 2).

Figure 2 Participant inclusion and exclusion procedure.

Although 40 datasets may seem limited at first glance, the structure of the experiment allows for a much larger effective sample size. Specifically, the 40 included participants collectively contributed approximately 480 minutes of mask-off (hypoxic) data. From these continuous recordings, we extract overlapping time windows to create individual data points for model training and testing. This approach results in a substantially expanded dataset, enabling meaningful statistical analysis and robust model development despite the relatively modest number of participants.

Preprocessing and feature extraction

In order to detect whether an aircrew member is hypoxic, a sequence of steps is undertaken that is well-established in scientific literature.

Preprocessing

Prior to the analysis, the data is preprocessed to remove unwanted artifacts that could negatively influence the model and prepare the data for further processing. See Figure 3 for an overview of the full pipeline.

Figure 3 Overview of the data preprocessing pipeline. ECG, electrocardiogram.

All ECG data was filtered using the SciPy Python library (46). First, the baseline wander was removed using an order 5 Butterworth highpass Infinite Impulse Response (IIR) filter (47) with a cutoff frequency of 0.5 Hz. Secondly, a filter was applied that removed the Power Line Inference (PLI) (48). This was done by smoothing the signal with a moving average kernel with the width of one period of 50 Hz (European power frequency).

Figure S4 shows an example of ECG data preprocessing. It displays raw ECG data and preprocessed ‘clean’ ECG data. The clean data has both baseline wander and powerline interference removed.

Segmentation

The data are split into smaller segments that entirely belong to a hypoxic section or to a nonhypoxic section of the dataset. The choice of the time interval of these segments is influenced by the need to perform HRV analysis on the ECG, the tools used [such as NeuroKit2 (49)], and the requirement for a sufficiently large dataset. The following constraints influence the choice of a time interval:

  • Minimum duration in literature: the minimum time interval for HRV analysis on ECG varies depending on specific conditions and analysis methods. A study (50) suggested that a longer measurement time is necessary for ultra-short-term HRV analysis under dynamic conditions, with at least 120 seconds required in post-exercise recovery or exercise conditions. Another study (51) further supported this, indicating that a minimum duration of 180 seconds may suffice for HRV measurement.
  • Technical limit: for segments shorter than 15 seconds, we encounter a technical limitation: not all HRV features computed from ECG with NeuroKit2 can be extracted, making this a hard lower-bound.
  • Data point count: The longer the segments, the fewer hypoxic data points are available for training and testing. As machine learning algorithms generally work better with more data points, shorter segments are preferred. In total, subjects are exposed approximately 480 minutes in a low oxygen environment with their mask-off (i.e., 12 minutes of total mask-off time for 40 subjects).

In order to be able to detect trends of HRV features over time, it is beneficial to provide data from multiple time windows to the classifier. In our case we choose three-time windows of different length. When splitting the data we make sure that the smallest time window, which is the most recent one, is either completely hypoxic or nonhypoxic. This windowing mechanism is visualized in Figure 4 for windows of 30, 60 and 120 seconds.

Figure 4 Data segmentation example: we show how data is divided in 30 seconds long data points and the extra portions of time used to extract descriptive features.

We call the entire data point hypoxic (with its three time windows), when the smallest window is hypoxic, even when the larger time windows cover nonhypoxic regions of the signal. As we are interested in detecting the early stages of hypoxia, such data points represent valuable transition periods.

Weighing the interests of having many data points and a robust HRV analysis, we have selected to compare two different splitting strategies:

  • 30-second hypoxic windows, with past windows of size 60 and 120 seconds, yielding 1.442 data points (half hypoxic, half non-hypoxic).
  • 20-second hypoxic windows, with past windows of 40 and 80 seconds, yielding 2.196 data points (half hypoxic, half non-hypoxic).

For each of these data points, the set of features described in Feature extraction are extracted for all three windows.

Feature extraction

In this section, we delve into the process of extracting features from the raw signals. The effectiveness of hypoxia classification heavily relies on the accurate and meaningful extraction of features. All described features can be applied to signals of variable length and are extracted for each of the three windows (see “Segmentation) and combined into a single feature vector. As an example, the Mean_ECG_Rate feature is extracted for the 30, 60 and 120 second time window, and added to the feature vector under the names 30S Mean_ECG_Rate, 60S Mean_ECG_Rate and 120S Mean_ECG_Rate. Similarly, for the 20-second dataset, measures such as Mean_ECG_Rate are extracted within the 20-second window and added to the feature vector as 20S Mean_ECG_Rate, and so forth. The features were extracted with the aid of the python library Neurokit2 (35) a package that is growing in popularity and is frequently cited in recent publications. Neurokit2 contains a total of 124 HRV measures that can be extracted from a signal. In our approach, we attempted to use all 124 Neurokit2 measures, but some measures were not computable for shorter time windows. In order to calculate the HRV features, a heartbeat detection was done on the ECG signal. This was done using Neurokit’s ecg process function which detects the R peaks in the ECG signal and then calculates the heart rate. The method used was elgendi2010. All available Neurokit2 measures where applied to all time windows, but some measures only worked on larger time windows and not the shorter ones. In this case, the Neurokit2 measure was omitted for the shorter time window, but still added for the longer time window. For the 20-second dataset a total of 86 Neurokit2 measures were usable for at least one time window, resulting in 218 features in total (Neurokit2 measure – time window combinations). For the 30-second dataset 90 Neurokit2 measures were usable, resulting in 253 features in total.

In order to be concise, we have chosen not to list all of the features that entered the feature selection process here, but only the Neurokit2 measures that were determined to be valuable for classifying hypoxia after feature selection (see “Feature selection methodology” and “Feature importance results”). Table S1 provides descriptions for these 18 Neurokit2 measures, as well as the three non-ECG measures.

Statistical analysis

A normality test was conducted on data from each feature. Both the Saphiro-Wilk test (52) and the D’Agostino test (53) returned a negative response, indicating the features do not follow a normal distribution. Therefore, we opted to normalize the features using Min-Max rescaling rather than a z-score standardization. This standardization was performed subject-wise, as that resulted in better performance.

Feature selection methodology

Feature selection is important as it helps to improve the model’s performance by eliminating irrelevant or redundant features, which can reduce overfitting and enhance generalization to new data. Additionally, it simplifies the model, making it more interpretable and reducing the computational cost of training and classification.

Feature selection may be performed by starting with a single feature and sequentially adding the next-most-valuable feature (forward selection), or by starting with all features and iteratively removing the least-valuable feature (backwards selection) (54-56). As there are many features available, we have opted for the forward selection method. At the beginning of the forward selection process, all available Neurokit2 measures (86/90, see “Feature extraction”), along with the mean RR, mean skin temperature, and standard deviation of skin temperature, are candidate features to be added.

The feature value is determined by a given quality measure using a given classification method, which is called a wrapper method (55). To gain a thorough understanding of what features are important to detect hypoxia, irrelevant of the classifier being used, we have chosen to evaluate five different classification methods and aggregate their results, of which parameters values where chosen based on preliminary experiments: (I) logistic regression, (II) K-nearest neighbors for k=5, (III) support vector classification (SVC), (IV) decision tree and (V) random forest1. These methods were chosen for their diversity, interpretability, and ability to provide confidence estimates. Unlike models such as one-dimensional convolutional neural networks (1D-CNNs) or dynamic time warping (DTW)—which are more suited for raw time-series data—our focus was on evaluating handcrafted features, for which these classifiers are more appropriate. The underlying model of each classifier impacts what features are most beneficial for the classification task. The value of a feature is judged by the change in accuracy of the classifier when the feature is included/excluded in the feature set.

We chose a cross-validation fold value of 5. Each of the five classifiers is used as a wrapper method, resulting in a total of five feature rankings (one for each classifier). The ranking score of the feature (5 points when being the most important, and 1 point for being the 5th most important feature) is accumulated across all feature rankings. The feature with the highest total score is deemed as the most important feature.

Hypoxia classification methodology

For each dataset (20, 30 seconds) we balanced the dataset with random under-sampling of the majority class (nonhypoxic) per subject. In this manner an equal amount of nonhypoxic and hypoxic datapoints per subject are present in the dataset, which helps the models to pay more attention to the hypoxic datapoints at the cost of possible discarding useful non-hypoxic information.

For the hypoxia classification task we measure the performance of the same five classifiers, and their earlier described settings, that were used during the feature selection (see “Feature selection methodology”). These models are trained using the subset of the five most relevant features for their classifier. To ensure an accurate estimation of the classifier’s performance, we employed a leave-one-subject-out cross-validation strategy, where data from one participant was used exclusively as the test set while the model was trained on data from the remaining participants. This approach guarantees that the training and test sets remain completely independent and prevents data leakage across participants. As performance metrics are averaged over all participants, this strategy also mitigates the impact of potential distributional differences between subsets, reducing the need for explicit comparisons of demographic or clinical variables across fixed data splits.

We assess model performance using two key metrics, namely the accuracy and the log loss.

The accuracy is straightforward to interpretate and widely accepted as a performance metric. Accuracy represents the proportion of correct predictions out of the total number of instances, providing a clear and easily understandable indication of the classifier’s effectiveness.

The log-loss performance measure is given by the equation:

LogLoss=1Ni=1N(yilog(pi)+(1yi)log(1pi))

where N is the total number of data points, yi represents the true label (0 or 1) of a datapoint i, pi is the predicted probability of the positive class for instance i. Log loss (also known as cross-entropy loss) quantifies the difference between predicted probabilities and actual labels. A lower log-loss value indicates a model that predicts probabilities closer to the true labels, reflecting better performance.

The log loss is a particularly useful measure for judging the performance of a classifier that provides confidence scores because it evaluates the accuracy of the predicted probabilities, not just the predicted classes. Unlike accuracy, which only considers whether the prediction was correct or not, Log Loss penalizes classifications that are confident but wrong more heavily. This ensures that the classifier is not only accurate but also calibrated, providing meaningful confidence scores that reflect the true likelihood of each class. By incorporating the probabilities, Log Loss offers a nuanced view of model performance, emphasizing the importance of producing reliable confidence estimates.

Additional metrics such as sensitivity, specificity, and area under the ROC curve (AUC) can offer an even more comprehensive evaluation of model performance and future work will incorporate these measures to further characterize classifier behavior.

The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the institutional ethics board of TNO (No. 2023-030) and informed consent was obtained from all individual participants.


Results

Feature importance results

The approach as described in “Feature selection methodology” is followed to determine what features are valuable for classifying hypoxia. This is done on the 20-second and on the 30-second dataset, to see whether other features are prominent for different time intervals. For a description of the features, we refer back to Table S1. Each feature is prefixed with the time window for which the Neurokit2 measure was analyzed. The features that are never chosen as top-5 feature by any classifier (and thus have a score of 0) are omitted from the figures.

Table 2 shows the most important features for the 20-second dataset. The most-important feature is the HRV_Prc80NN on the 40-second time window, followed by the 80S temperature_mean.

Table 2

Cross model feature ranking for the 20-second dataset (5-fold cross validation)

Feature Score
40S HRV_Prc80NN 11
80S temperature_mean 7
80S HRV_HFn 5
80S HRV_MinNN 5
80S ECG_Rate_Mean 5
80S HRV_Prc80NN 5
80S HRV_MFDFA_alpha1_Asymmetry 5
80S HRV_pNN20 4
80S HRV_LZC 4
40S resp_mean 4
40S HRV_GI 3
80S resp_mean 3
40S temperature_mean 3
20S HRVpNN20 2
20S HRV_MeanNN 2
40S HRV_MaxNN 2
40S HRV_MFDFA_alpha1_Fluctuation 2
20S temperature_mean 1
20S temperature_std 1
40S temperature_std 1

The score is the sum of the score per classifier. A feature can have a maximum score of 25 if it was the top-1 feature for all 5 classifiers.

Table 3 shows the same analysis, but now on the 30-second dataset. Here, the 120S resp_mean jumps out as most important feature. This feature was also deemed important on the 20-second dataset Table 2, although somewhat less important.

Table 3

Cross model feature ranking for the 30-second dataset (5-fold cross validation)

Feature Score
120S resp_mean 20
60S HRV_Prc80NN 6
120S HRV_MCVNN 5
120S HRV_Prc80NN 5
120S temperature_std 5
60S HRV_MadNN 4
120S temperature_mean 4
120S HRV_HFD 4
60 HRV_Prc20NN 4
30S ECG_rate_mean 4
60S HRV_MeanNN 3
60S temperature_std 2
60S HRV_MedianNN 2
60S HRV_MadNN 2
30S temperature_mean 2
30S HRV_PSS 1
30S HRV_PIP 1
30S HRV_SD2d 1

The score is the sum of the score per classifier. A feature can have a maximum score of 25 if it was the top-1 feature for all 5 classifiers.

Hypoxia classification results

In this section, we compare model performances on both the 20- and 30-second hypoxia window datasets. The accuracy and log loss performance of each classifier, including the selected features for each classifier, are provided for the 20-second time windows in Table 4 and 30-second time windows in Table 52.

Table 4

The performance of five classifiers on the 20-second dataset (leave-one-out cross validation)

Classifier Accuracy Log loss Top features
SVC 0.90±0.07* 0.25±0.15* 80S HRV_MinNN
40S HRV_Prc80NN
40S HRV_GI
40S temperature_mean
40S resp_mean
Decision tree 0.90±0.07* 0.72±1.43 80S HRV_Prc80NN
80S HRV_LZC
80S resp_mean
40S HRV_MaxNN
40S temperature_mean
Random forest 0.89±0.07 0.32±0.10 80S HRV_MFDFA_alpha1_Asymmetry
80S temperature_mean
40S HRV_Prc80NN
40S HRV_MFDFA_alpha1_Fluctuation
40S temperature_std
Logistic regression 0.88±0.07 0.30±0.16 80S HRV_HFn
40S HRV_Prc80NN
40S resp_mean
20S HRV_pNN20
20S temperature_mean
KNN 0.86±0.07 1.17±1.07 80S ECG_Rate_Mean
80S HRV_pNN20
80S temperature_mean
20S HRV_MeanNN
20S temperature_std

Data are presented as mean ± standard deviation. *, the best performance values. KNN, K-nearest neighbors; SVC, support vector classification.

Table 5

The performance of five classifiers on the 30-second dataset (leave-one-out cross validation)

Classifier Accuracy Log loss Top features
SVC 0.89±0.08 0.27±0.17* 120S HRV_MCVNN
120S temperature_mean
120S resp_mean
30S ECG Rate_Mean
30S HRV_PSS
Decision tree 0.90±0.07* 0.55±0.58 120S temperature_std
120S resp_mean
60S HRV_MeanNN
60S HRV_MedianNN
30S HRV_SD2d
Random forest 0.88±0.08 0.32±0.11 120S resp_mean
60S HRV_Prc20NN
60S HRV_Prc80NN
30S HRV_MadNN
30S temperature_mean
Logistic regression 0.89±0.08 0.31±0.18 120S resp_mean
60S HRV_MadNN
60S HRV_Prc80NN
30S ECG_Rate_Mean
30S temperature_mean
KNN 0.85±0.08 1.19±1.33 120S HRV_Prc80NN
120S HRV_HFD
120S resp_mean
60S temperature_std
30S HRV_PIP

Data are presented as mean ± standard deviation. *, the best performance values. KNN, K-nearest neighbors; SVC, support vector classification.

To understand what type of mistakes a model makes, we graphically inspect the classified data points using the SVC model. Figure 5 shows all data points for the 20 s time windows for all participants. In the top graph the true positives (green) and false negatives (red) are shown. All data points are on the right side of the yellow line, as in that area of the experiment the aircrew members were exposed to hypoxic conditions. The bottom graph shows true negatives (green) and false positives (red). For both graphs the y-axis is the models confidence for its classification.

Figure 5 The classification of data points of the SVC classifier on the 20 s dataset. Top: hypoxic (positive class) data points as per ground truth label, red dots are mislabeld data points (false negatives). Bottom: nonhypoxic (negative class) data points as per ground truth label, red dots are mislabeled data points (false positives). The yellow line marks where the rapid decompression starts, which is when the mask-off events are about to begin. The y-axis represents the confidence of the model. FP, false positive; SVC, support vector classification; TN, true negative.

In order to better understand the workings of the classifier, we now investigate a similar graph of all data points, but now with the most important feature for the SVC model in the 20 s dataset as y-axis label. Figure 6 shows this adapted plot. Please note that the coloring and x-axis has not changed to Figure 5 only the height of the points.

Figure 6 The classification of data points of the SVC classifier on the 20 s dataset. Top: hypoxic (positive class) data points as per ground truth label, red dots are mislabeled mislabeled data points (false negatives). Bottom: nonhypoxic (negative class) data points as per ground truth label, red dots are mislabeled data points (false positives). The yellow line marks where the rapid decompression starts, which is when the mask-off events are about to begin. The y-axis represents the normalized value of the 80S HRV_MinNN feature. FN, false negative; SVC, support vector classification; TP, true positive.

Discussion

Key findings

The goal of this study was to investigate using ECG and temperature signals to detect hypoxic conditions of military aircrew members due to its non-invasive nature, ease of integration into existing flight equipment, and ability to provide continuous real-time data. A unique dataset was created by measuring 40 military aircrew members with multiple physiological and physical data sensors undergoing hypoxia training in a hypobaric chamber. During the hypoxia training, a standardized flight protocol was simulated with different altitudes (by changing air pressure), including rapid decompression and the temporary removal of the trainee’s oxygen masks. Besides the air pressure and manually collected mask on/off events, personal sensors captured the trainee’s skin temperature, PPG, ECG, RR and SpO2. As part of data quality curation, data from three trainees were removed due to invalid ECG measurements and from two trainees due to unreliable SpO2 measurements, resulting in 40 valid signals. Taken together, this resulted in a unique curated multi-subject, multi-sensor, time-series dataset with induced and measured hypoxia.

The oxygen mask being off (labeled as “hypoxic”) or being on (labeled as “nonhypoxic”) at high altitudes, was chosen as the ground truth label to train the classification algorithms on. Unfortunately, we could not reliably label hypoxia based on SpO2 measurements as SpO2 measurements below 95% occurred too frequently at moments where hypoxia was not possible.

Features were generated according to a 20-second data point approach (using features with window-sizes of 20, 40 and 80 seconds) as well as a 30-second data point approach (generating features with window-sizes of 30, 60 and 120 seconds). For the 20-second dataset, a total of 218 features were tested, and for the 30-second dataset 253 features were tested in total.

A feature selection approach was implemented that scored features based on their occurrence in the best-performing feature rankings across five different classifiers. When comparing the results for the 20- and 30-second datasets (Tables 2,3), we observe that several features—HRV_Prc80NN, temperature_mean, resp_mean, ECG_rate_Mean, and temperature_std—consistently appear across various time windows and stand out as promising candidates for classifying hypoxia. This selection approach also revealed a clear preference for features derived from longer time windows. Specifically, features originating from the smallest windows (20 s in Table 2 and 30 s in Table 3) are considered valuable by some classifiers but appear less frequently than those from extended windows (e.g., 60, 80 and 120 s). These findings suggest that longer ECG segments provide more reliable indicators of hypoxia and offer valuable guidance for optimizing time window selection in future model development.

For the classification task (Tables 4,5), differences in classifier performance are not very large, suggesting that the type of classifier is not the most critical component for achieving high performance. Among the five classifiers tested, SVC performed best overall—achieving top-1 and top-2 rankings in accuracy for the 30- and 20-second datasets, respectively—and consistently produced the lowest log loss. This indicates more confident and well-calibrated predictions. Such performance aligns with the nature of the hypoxia classification problem, which is likely non-linearly separable due to complex physiological signal interactions; SVC is particularly suited for this task due to its ability to define flexible decision boundaries. In contrast, Decision Trees, despite occasionally achieving high accuracy, showed substantially higher log loss, reflecting poor probability calibration and a tendency to overfit the training data. Random forests and logistic regression provided a good balance across performance metrics, while K-nearest neighbors lagged behind, likely due to its sensitivity to local noise and limitations in high-dimensional feature spaces. Importantly, features like resp_mean and temperature_mean consistently appeared across models, highlighting their relevance in detecting hypoxic states. The variation in additional features selected by each model further emphasizes their differing inductive biases and sensitivity to specific signal characteristics. Taking into account that all classifiers yielded high and balanced performance across diverse metrics, we conclude that the models achieve robust and reliable hypoxia classification based on the provided features.

We found that the longer time windows were preferred by the classification models. This leads to the question whether even larger time windows are also valuable for the classification task. We would like to extend our study towards larger time windows in the future. Special attention, however, has then to be given towards the consequential reduction of the number of data points when the time windows get longer.

Deeper analysis of the commonality of the current misclassifications may lead to further improvements. In the current study a deep investigation under what conditions the classifiers are reliable was out of scope. The results of such an analysis may generate valuable insights that lead to further improvements of the classifier or to clear limitations on the usage of the models.

Strength and limitations

The current study population primarily consists of fit military aircrew members (7 female, 33 male). To ensure our model’s generalizability to a broader population, it is important to account for the variability in physiological responses to hypoxia among different individuals. The model’s performance may be affected by specific physiological characteristics of less fit military aircrew members, particularly in response to varying environmental conditions encountered during actual flights. To enhance and validate the model’s generalizability, it is essential to incorporate a more diverse population into the training dataset. Furthermore, continuous validation and refinement of the model with new data from different populations and real-world scenarios will improve its robustness and reliability.

The data was collected in a hypobaric chamber. The aircrew members were sitting still and performing their training. When an aircrew member is in flight, the situation significantly changes. There are high G-forces, and vibrations, leading to additional noise in the sensor readings. The wearing of additional gear may also impact the sensor readings. The aircrew member may also be preoccupied with the mission, rather than paying attention to hypoxia symptoms.

It would be valuable to collect data from aircrew members during flight. However, exposing the aircrew members to hypoxic events during flight is dangerous and not ethical. We therefore should investigate how models for hypoxia detection in safe conditions can transfer towards a use in a real fighter jet.

We have chosen the manual labels of when aircrew members take of their mask instead of the SpO2 sensor data. Although related, this is not an equal substitute. Specifically for the transition periods between non-hypoxic and hypoxic and vice versa, is the mask on/off label not an indication of hypoxia itself. The trained classifier is essentially a classifier whether the aircrew member is exposed to hypoxic conditions, rather than that the aircrew member is hypoxic. The benefit of this approach is that early stages of hypoxia may be detected, and the aircrew member may intervene before becoming hypoxic.

Comparison with similar research

Previous studies have explored hypoxia classification using machine learning, particularly through ECG and HRV features. Oliveira et al. (32) and Castro-Herrera et al. (33) demonstrated the potential of HRV metrics to reflect autonomic responses to hypoxia in healthy individuals and pilots, respectively, while Aebi et al. (34) investigated hypobaric influences in pilot trainees. In fetal hypoxia detection, Dong et al. (37) and Hussain et al. (38) utilized SVMs and HRV features, and Ma’sum et al. (39) employed deep learning directly on ECG signals. For adult populations, Niaussat et al. (40) applied XGBoost and random forests to ICU patients using a small set of physiological parameters. Compared to these works, our study distinguishes itself through its focus on healthy military aircrew, the combination of ECG and skin temperature features, and the use of a diverse and curated multi-sensor dataset collected in a controlled hypobaric environment. Moreover, our feature selection strategy across multiple classifiers and time windows further refines the reliability of hypoxia detection.

Explanations of findings

When investigating the performance of the SVC, we observe that most misclassifications occur when the model indicates that it has low confidence in its prediction. This shows that the confidence estimations have been learned successfully by the model, and the confidence estimation values may be used in the resulting system. For example, by only warning an aircrew member when the model detects hypoxia and the confidence of the model for that prediction is high.

In general, we are satisfied to assess that most of the false classifications are concentrated in the second half of the experiment session (see Figure 5), essentially where the mask-off events and the depressurization of the chamber takes place. In the first part of the session only few data points are wrongly marked as hypoxic, while having low confidence in the classification. When regarding false negatives, we observe that the majority is happening at the start of the mask-off event. This makes sense, as the aircrew member has just taken of the mask and is not oxygen deprived yet. We observe that the classifier quickly detects when the pilot is exposed to hypoxic conditions after the mask is removed. Such a classifier could potentially warn aircrew members before serious symptoms manifest. When working with confidence levels, the general hope is that all misclassifications belong to data points for which also a low classification confidence is given. In Figure 5 we observe that this is true for the false negatives, but less so for false positives.

When regarding hypoxic data points (the top figure of Figure 6), we observe that the model makes mistakes when the 80S HRV_MinNN feature is having high or low values. When this feature has a value in the middle of the observed data range, the model made no mistakes. There is no clear pattern visible for the false negatives. Such an analysis can be made for all available features in order to better understand under what circumstances the classifier makes mistakes. This may lead to an even better model for the classification task, and we leave this as future research.

Implications and actions needed

To advance this research towards practical implementation, we propose the following directions:

  • Validation under gradual hypoxia: conduct follow-up studies with gradual hypoxia onset to better mirror real in-flight conditions and improve classifier generalization.
  • Sensor integration: it would be beneficial to use a more reliable sensor for detecting the true hypoxic state of aircrew members. In the future, a system may be envisioned that provides staged feedback to the aircrew member (e.g., you are at risk of becoming hypoxic, you are hypoxic). The fact that hypoxia was difficult to reliably capture hypoxia with the SpO2 sensor in our experiment, strengthens the need of our more practical approach. A good starting point are the various sensors that the Health Patch contains (44).
  • Model deployment & feedback modalities: investigate real-time system integration, including feedback modalities (auditory, visual, haptic) and user-centered alert mechanisms for in-flight use.
  • Expanded physiological event detection: extend classification to other events (e.g., arrhythmia, CBRN exposure) using additional biosignals from the Health Patch.
  • Population diversity: explore gender-specific physiological responses and inter-subject variability to enhance model robustness.
  • Reusable ML platform: continue developing our digital platform to streamline sensor data collection, enable cross-study learning, and simplify workflow reuse for both technical and non-technical users. See (57) for more information about this platform.
  • Application to other domains: explore the use in clinical and high-risk environments such as ICUs (40), fetal hypoxia monitoring (37-39).

Conclusions

This study describes the development of an algorithm for classifying hypoxic conditions of military aircrew members. We have shown that hypoxia can be predicted in a realistic setting from an ECG and skin temperature sensor, which opens the way to advance this approach towards practical applications to support military aircrew members in the cockpit.


Acknowledgments

We would like to thank QM_Project, and specifically Matteo Bigogno, for the collaboration on developing the HPMA platform that can run the developed hypoxia classifier. Without the support of the CMA operators, the data collection of aircrew members inside the hypobaric chamber would not have been possible. We furthermore thank 2M Engineering for their support during the use of their Health Patch sensor to collect the aircrews physiological data.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2024-469/rc

Data Sharing Statement: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2024-469/dss

Peer Review File: Available at https://jmai.amegroups.com/article/view/10.21037/jmai-2024-469/prf

Funding: This study was funded and conducted within the MEA IP project ‘HPMA’.

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jmai.amegroups.com/article/view/10.21037/jmai-2024-469/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the institutional ethics board of TNO (No. 2023-030) and informed consent was obtained from all individual participants.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.

1We set the maximum depth to 4 for both decision tree and random forest models to avoid overfitting.

2The feature rankings of these tables form the basis of the data of Table 2 and Table 3. Please note however that the accuracy and Log Loss measures are based on a leave-one-out cross-validation methodology, while the feature selection used a 5-fold cross validation methodology.


References

  1. Hall JE, Hall ME. Guyton and Hall Textbook of Medical Physiology. 14th ed. Philadelphia, PA: Elsevier; 2020.
  2. West JB, Schoene R, Luks A, et al. High Altitude Medicine and Physiology. 5th ed. Boca Raton, FL: CRC Press; 2012.
  3. Gradwell D, Rainford DJ. Ernsting’s Aviation and Space Medicine. 6th ed. Boca Raton, FL: CRC Press; 2024.
  4. Burtscher J, Niedermeier M, Hüfner K, et al. The interplay of hypoxic and mental stress: Implications for anxiety and depressive disorders. Neurosci Biobehav Rev 2022;138:104718. [Crossref] [PubMed]
  5. Kourtidou-Papadeli C, Papadelis C, Koutsonikolas D, et al. High altitude cognitive performance and COPD interaction. Hippokratia 2008;12:84-90.
  6. Neuhaus C, Hinkelbein J. Cognitive responses to hypobaric hypoxia: implications for aviation training. Psychol Res Behav Manag 2014;7:297-302. [Crossref] [PubMed]
  7. Pun M, Guadagni V, Bettauer KM, et al. Effects on Cognitive Functioning of Acute, Subacute and Repeated Exposures to High Altitude. Front Physiol 2018;9:1131. [Crossref] [PubMed]
  8. Cable GG. In-flight hypoxia incidents in military aircraft: causes and implications for training. Aviat Space Environ Med 2003;74:169-72.
  9. Blacker KJ, McHail DG. Effects of Acute Hypoxia on Early Visual and Auditory Evoked Potentials. Front Neurosci 2022;16:846001. [Crossref] [PubMed]
  10. Cobb AB, Levett DZH, Mitchell K, et al. Physiological responses during ascent to high altitude and the incidence of acute mountain sickness. Physiol Rep 2021;9:e14809. [Crossref] [PubMed]
  11. Steinman Y, Groen E, Frings-Dresen MHW. Hypoxia impairs reaction time but not response accuracy in a visual choice reaction task. Appl Ergon 2023;113:104079. [Crossref] [PubMed]
  12. Wang Y, Yu X, Liu Z, et al. Influence of hypobaric hypoxic conditions on ocular structure and biological function at high attitudes: a narrative review. Front Neurosci 2023;17:1149664. [Crossref] [PubMed]
  13. Ortiz-Prado E, Dunn JF, Vasconez J, et al. Partial pressure of oxygen in the human body: a general review. Am J Blood Res 2019;9:1-14.
  14. Alvear-Catalán M, Montiglio C, Aravena-Nazif D, et al. Oxygen Saturation Curve Analysis in 2298 Hypoxia Awareness Training Tests of Military Aircrew Members in a Hypobaric Chamber. Sensors (Basel) 2024;24:4168. [Crossref] [PubMed]
  15. Chan ED, Chan MM, Chan MM. Pulse oximetry: understanding its basic principles facilitates appreciation of its limitations. Respir Med 2013;107:789-99. [Crossref] [PubMed]
  16. Louie A, Feiner JR, Bickler PE, et al. Four Types of Pulse Oximeters Accurately Detect Hypoxia during Low Perfusion and Motion. Anesthesiology 2018;128:520-30. [Crossref] [PubMed]
  17. Wick KD, Matthay MA, Ware LB. Pulse oximetry for the diagnosis and management of acute respiratory distress syndrome. Lancet Respir Med 2022;10:1086-98. [Crossref] [PubMed]
  18. Martin D, Johns C, Sorrell L, et al. Effect of skin tone on the accuracy of the estimation of arterial oxygen saturation by pulse oximetry: a systematic review. Br J Anaesth 2024;132:945-56. [Crossref] [PubMed]
  19. Luks AM, Swenson ER. Pulse oximetry at high altitude. High Alt Med Biol 2011;12:109-19. [Crossref] [PubMed]
  20. Akkermans MA, Sillen MJ, Wouters EF, et al. Validation of the oxycon mobile metabolic system in healthy subjects. J Sports Sci Med 2012;11:182-3.
  21. Salier Eriksson J, Rosdahl H, Schantz P. Validity of the Oxycon Mobile metabolic system under field measuring conditions. Eur J Appl Physiol 2012;112:345-55. [Crossref] [PubMed]
  22. Wasserman K, Hansen JE, Sue DY, et al. Principles of Exercise Testing and Interpretation. 5th ed. Amsterdam, The Netherlands: Lippincott Williams & Wilkins; 2011.
  23. McClung HL, Tharion WJ, Walker LA, et al. Using a Contemporary Portable Metabolic Gas Exchange System for Assessing Energy Expenditure: A Validity and Reliability Study. Sensors (Basel) 2023;23:2472. [Crossref] [PubMed]
  24. Wei Y, Jin L, Wang S, et al. Hypoxia detection for confined-space workers: Photoplethysmography and machine-learning techniques. SN Comput Sci 2022;3:290.
  25. Sawangjai P, Seesawad N, Wilaiprasitporn T. Removal of Motion Artifacts From the PPG Signal Using Attentive Generative Adversarial Networks With Dual Discriminator. IEEE Trans Instrum Meas 2025;74:2504210.
  26. Fine J, Branan KL, Rodriguez AJ, et al. Sources of Inaccuracy in Photoplethysmography for Continuous Cardiovascular Monitoring. Biosensors (Basel) 2021;11:126. [Crossref] [PubMed]
  27. Moerman A, Wouters P. Near-infrared spectroscopy (NIRS) monitoring in contemporary anesthesia and critical care. Acta Anaesthesiol Belg 2010;61:185-94.
  28. Perrey S, Quaresima V, Ferrari M. Muscle Oximetry in Sports Science: An Updated Systematic Review. Sports Med 2024;54:975-96. [Crossref] [PubMed]
  29. Nonin Medical Inc. Nonin Model 8000R Reflectance Pulse Oximeter Sensor: A Non-Invasive Technology for Measuring Tissue Oxygenation. 2011. Available online: https://www.nonin.com. Accessed: 2024-11-06.
  30. Pigat L, Geisler BP, Sheikhalishahi S, et al. Predicting Hypoxia Using Machine Learning: Systematic Review. JMIR Med Inform 2024;12:e50642. [Crossref] [PubMed]
  31. Herzig JJ, Ulrich S, Schneider SR, et al. Heart rate variability in pulmonary vascular disease at altitude: a randomised trial. ERJ Open Res 2024;10:00235-2024. [Crossref] [PubMed]
  32. Oliveira ALMB, Rohan Pde A, Gonçalves TR, et al. Effects of hypoxia on heart rate variability in healthy individuals: A systematic review. Int J Cardiovasc Sci 2017;30:251-61.
  33. Castro-Herrera JM, Cedeño-Serna JC, Tuta-Quintero E, et al. Heart rate variability as a predictor of hypobaric hypoxia in aircraft pilots. Revista Latinoamericana de Hipertensión 2021;314-20.
  34. Aebi MR, Bourdillon N, Bron D, et al. Minimal Influence of Hypobaria on Heart Rate Variability in Hypoxia and Normoxia. Front Physiol 2020;11:1072. [Crossref] [PubMed]
  35. Frasch MG. Comprehensive HRV estimation pipeline in Python using Neurokit2: Application to sleep physiology. MethodsX 2022;9:101782. [Crossref] [PubMed]
  36. Pham T, Lau ZJ, Chen SHA, et al. Heart Rate Variability in Psychology: A Review of HRV Indices and an Analysis Tutorial. Sensors (Basel) 2021;21:3998. [Crossref] [PubMed]
  37. Dong S, Boashash B, Azemi G, et al. Automated detection of perinatal hypoxia using time-frequency-based heart rate variability features. Med Biol Eng Comput 2014;52:183-91. [Crossref] [PubMed]
  38. Hussain NM, Amin B, McDermott BJ, et al. Feasibility Analysis of ECG-Based pH Estimation for Asphyxia Detection in Neonates. Sensors (Basel) 2024;24:3357. [Crossref] [PubMed]
  39. Ma’sum MA, Intan PRD, Jatmiko W, et al. Improving deep learning classifier for fetus hypoxia detection in cardiotocography signal. In: 2019 International Workshop on Big Data and Information Security (IWBIS); 2019; Piscataway, NJ. IEEE; 2019. p. 51-56. doi:10.1109/IWBIS.2019.8935835.
  40. Niaussat V, Giguère R, Paul TS, et al. Real-time machine learning for ICU hypoxia prediction: A pilot study. Neuroergonomics and Cognitive Engineering. 2024;126:12-22.
  41. Zhang X, Zhang Y, Si Y, et al. A high altitude respiration and SpO2 dataset for assessing the human response to hypoxia. Sci Data 2024;11:248. [Crossref] [PubMed]
  42. Jain KK. Hyperbaric Chambers: Equipment, Technique, and Safety. In: Textbook of Hyperbaric Medicine. Cham: Springer; 2017. p. 61-78.
  43. Ministerie van Defensie. Drukkamer. 2024. Available online: https://www.defensie.nl/binaries/content/gallery/defensie/content-afbeeldingen/organisatie/luchtmacht/centrum-voor-mens-en-luchtvaart/drukkamer004.jpg. Accessed: 2024-08-02.
  44. 2M Engineering. Draagbaar Gezondheidspleister. 2024. Available online: https://www.2mel.nl/nl/draagbare-gezondheidspatch/. Accessed: 2024-11-13.
  45. Grimnes S, Martinsen ØG. Bioimpedance and Bioelectricity Basics. 3rd ed. London: Academic Press; 2015. doi:10.1016/C2012-0-06951-7.
  46. SciPy. SciPy Documentation. 2024. Available online: https://scipy.org/ (Accessed: 2024-11-14).
  47. Ozaydin S, Ahmad I. Comparative Performance Analysis of Filtering Methods for Removing Baseline Wander Noise from an ECG Signal. Fluct Noise Lett. 2024;23:2450046.
  48. Mian Qaisar S. Baseline wander and power-line interference elimination of ECG signals using efficient signal-piloted filtering. Healthc Technol Lett 2020;7:114-8. [Crossref] [PubMed]
  49. École de Neuropsychologie. NeuroKit2 Documentation. 2024. Available online: https://github.com/neuropsychology/NeuroKit. Accessed: 2024-11-14.
  50. Kim JW, Seok HS, Shin H. Is Ultra-Short-Term Heart Rate Variability Valid in Non-static Conditions? Front Physiol 2021;12:596060. [Crossref] [PubMed]
  51. Choi WJ, Lee BC, Jeong KS, et al. Minimum Measurement Time Affecting the Reliability of the Heart Rate Variability Analysis. Korean J Health Promot 2017;17:269-74.
  52. Noughabi HA. Two Powerful Tests for Normality. Ann Data Sci 2016;3:225-34.
  53. Krouse DP. The power of D’Agostino’s D test of normality against a normal mixture alternative. Commun Stat Theory Methods 1994;23:47-57.
  54. Glaysher P, Katzy JM, An S. Iterative subtraction method for Feature Ranking. 2019. Available online: https://ar5iv.labs.arxiv.org/html/1906.05718
  55. Jović A, Brkić K, Bogunović N. A review of feature selection methods with applications. In: 2015 38th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO); 2015; p. 1200-5.
  56. O'Hara S, Wang K, Slayden RA, et al. Iterative feature removal yields highly discriminative pathways. BMC Genomics 2013;14:832. [Crossref] [PubMed]
  57. TNO. The Human Performance Monitoring Algorithms project (HPMA): a software module to interpret sensor data of human performance monitoring sensors. 2024. Available online: https://diamonds.tno.nl/projects/hpma. Accessed: 2025-2-3.
doi: 10.21037/jmai-2024-469
Cite this article as: Schadd MPD, van Baardewijk JU, Bojczuk MT, Aliberti A, Vuijk FL, Linssen L, Gijsbertse K, Houben MMJ, Arrigoni-Neri M, Kingma BRM, van Someren EP. Model development and validation for classifying hypoxia in military aircrew using ECG and skin temperature. J Med Artif Intell 2026;9:5.

Download Citation