Figures
Abstract
Existing forest fire risk prediction methods often focus on single-element analysis, which makes it difficult to effectively capture the underlying mechanisms of the “fuel-climate” interaction. This paper proposes a physically aligned prediction framework named FWI-MSNet. Using 18 years of synchronized observation data from the Huitong Ecological Station in China, the framework constructs a multi-scale feature system, selecting 21 physically relevant key features, including core indicators of the Forest Fire Weather Index (FWI). A parallel multiscale 1-Dimensional Convolutional Neural Network (1D-CNN) extracts dynamic features across multiple temporal scales (from daily to seasonal). Subsequently, a synergistic mechanism integrating Gated Recurrent Units (GRU) and Transformer achieves deep integration of temporal evolution and global correlation. Experimental results demonstrate that compared to seven baseline models (including XGBoost, LSTM, CNN, and four ablation variants), the mean of RMSE, MAE, and MAPE decreases by 56.1%, and R2 improves from 0.223 (baseline average) to 0.9251. Furthermore, it accurately captures the trajectory of fire risk evolution in the 2020 catastrophic forest fires in Australia. This study provides a valuable reference for forest fire risk prediction.
Citation: Chen B, Zeng A, Xie Y, Wang X, Zhu L, Xie Q, et al. (2026) Physically aligned forest fire risk prediction: A deep learning framework coupling fuel and climate multivariate factors. PLoS One 21(9): e0355829. https://doi.org/10.1371/journal.pone.0355829
Editor: Julia A. Jones, Oregon State University, UNITED STATES OF AMERICA
Received: January 6, 2026; Accepted: July 24, 2026; Published: September 1, 2026
Copyright: © 2026 Chen et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The complete dataset is derived from three datasets provided by the National Research Station of Huitong Forest Ecosystems (2005-2022), available through the China Ecological Science Data Center (https://www.nesdc.org.cn/) under DOIs: 10.12199/nesdc.ecodb.2021YFF0703900.htf.2025.1, 10.12199/nesdc.ecodb.2021YFF0703900.htf.2025.2, and 10.12199/nesdc.ecodb.2021YFF0703900.htf.2025.3. A demo dataset and the complete source code are available at the GitHub repository: https://github.com/CBJYB/FWIMSNet-Forest-Fire-Risk-Prediction-Framework. The ERA5 reanalysis data used for the case study are openly available from the Copernicus Climate Change Service: https://doi.org/10.24381/cds.adbb2d47.
Funding: This study was supported by the Yibin University Research Project (Grant No. 412-2024XJPY06). The funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. No additional external funding was received for this study.
Competing interests: The authors have declared that no competing interests exist.
1. Introduction
Forest fires, driven by the complex interplay of meteorological conditions, fuel characteristics, and anthropogenic activities [1], have become a major global concern. As climate change intensifies, the frequency, intensity, and duration of the wildfire season have risen significantly, posing severe threats to ecological balance, human safety, and socioeconomic development [2–4]. Data from the China Statistical Yearbook (2000–2023) corroborate this trend, revealing notable fluctuations and periodic peaks in key fire indicators (Fig 1).
From 2000 to 2023, the frequency of forest fires and the direct economic losses they caused in China showed an upward trend, while the area of forest affected decreased significantly. This trend reflects that, although the burned area per single fire has been effectively controlled, the increasing frequency of fires and the associated economic losses have not diminished correspondingly, suggesting that fire risk may be more influenced by the interactive effects of fuel accumulation and climatic conditions.
Traditional early warning systems primarily rely on physical sensor technologies within Wireless Sensor Networks (WSNs), utilizing temperature, humidity, and smoke detectors for real-time monitoring [5–7]. However, the large-scale deployment of WSNs presents persistent challenges related to energy consumption, network lifespan, and real-time responsiveness, which limit their predictive capability for large-area risk assessment [8].
Consequently, researchers have increasingly turned to data-driven approaches. Among these, machine learning and multi-source data fusion have driven recent progress in forest fire risk assessment. Abbas Khurram [9] et al. used random forest and gradient boosting machines to confirm temperature and relative humidity as critical factors. Jiajun Chen [10] et al. built a risk model using exponentially weighted decay and support vector machine regression, achieving high test accuracy. Diana Škurić Kuraži [11] et al. proposed a new soil-based fire risk index that outperforms traditional weather indices. Although these methods demonstrate utility, their predictive accuracy is often constrained by a limited capacity to model complex nonlinear spatiotemporal dependencies.
In contrast, deep learning approaches have been increasingly adopted for forest fire risk prediction due to their ability to model complex spatiotemporal patterns. Recent comparative studies have also evaluated the performance of CNN, LSTM, GRU, and Transformer architectures for time series forecasting in various domains [12]. Mohammad Marjani [13] et al. developed a CNN-BiLSTM hybrid model for near-real-time wildfire spread prediction, achieving an Intersection over Union (IoU) of 0.58 and an F1 score of 0.73, demonstrating the effectiveness of combining convolutional feature extraction with sequential modeling. Xufeng Lin [14] et al. proposed a forest fire prediction model based on LSTNet, which captures both short-term and long-term temporal dependencies; the model achieved a high accuracy of 0.941, highlighting the importance of utilizing spatial background information and the periodic nature of fire-related factors. More recently, Xinyu Miao [15] et al. introduced a window-enhanced Transformer model for time series prediction, achieving strong performance (ACC = 91.56%, RMSE = 0.37, MAE = 0.05) by effectively leveraging spatial context and periodic patterns in forest fire drivers. Despite these advances, most deep learning models treat fire prediction as a purely data-driven task, with limited integration of the physical principles governing fire behavior—such as the multi-scale temporal dynamics embedded in the Fire Weather Index (FWI) system.
Furthermore, researchers have attempted to enhance physical interpretability by improving traditional indices. Santos L C [16] et al. developed an enhanced FWI (FWIe), which correlates indices with satellite-detected fires via fire radiative power, thereby improving risk assessment. Binita Kumari [17] et al. evaluated risk using various triggering factors, finding that GIS-based risk areas closely match historical fire occurrences. These efforts highlight the growing recognition of integrating physical knowledge with data-driven methods. However, most existing approaches either treat physical indices as mere input variables without exploiting their inherent multi-scale temporal structure, or rely on static spatial risk mapping that fails to capture the dynamic coupling between fuel and climate factors across different time scales.
Despite these significant efforts, several critical research gaps remain. First, most existing deep learning models [18] treat fire prediction as a purely data-driven task, often overlooking the inherent multi-scale temporal structure of the physical processes governing fire behavior [19] (e.g., daily weather-driven responses vs. seasonal soil moisture accumulation). Second, although the Fire Weather Index (FWI) system [20] provides a well-established physical framework for characterizing fire danger, its potential to serve as a physically aligned organizational principle for structuring multi-source features—rather than merely an additional input variable—remains largely untapped [21]. Third, existing architectures struggle to simultaneously capture the multi-scale temporal patterns [22] (from daily to seasonal variations) and long-term dependencies inherent in the coupled fuel-climate system [23], limiting their ability to align these interacting physical processes.
To bridge these gaps, this paper proposes FWI-MSNet, a novel physically aligned deep learning framework for the synchronous analysis of fuel-climate coupling. This study uses empirical data from the China Ecological Science Data Center (https://www.nesdc.org.cn/). The novelty of this framework lies in three aspects: physically aligned feature organization based on the FWI system‘s multi-scale temporal structure; explicit physical mapping between CNN kernel sizes and FWI component time scales; and a GRU-Transformer collaborative mechanism for temporal-global fusion. These design choices translate into three main contributions:
- Physically Aligned Feature Organization: Unlike conventional black-box models that treat all inputs uniformly, our framework uses the Fire Weather Index (FWI) system as a physically meaningful basis for feature engineering. By organizing input features according to the FWI’s inherent multi-scale structure (short-term, medium-term, long-term), the learning process aligns with the physical principles of fire dynamics, leading to more consistent and interpretable predictions.
- Multi-Scale Feature Extraction: We design parallel multi-scale 1D-CNN branches with kernel sizes aligned to the physical timescales of fire risk formation (daily, monthly, and quarterly). This enables the model to capture the diverse temporal patterns inherent in meteorological and fuel data—from short-term weather-driven responses to long-term drought accumulation—ensuring that features across all relevant physical scales are represented for subsequent fusion.
- Sequential Modeling with Temporal-Global Fusion: We integrate a GRU-Transformer hybrid mechanism to model both the temporal evolution and global contextual dependencies within the multi-scale features. The GRU layer captures mid-term sequential patterns, while the Transformer encoder extracts cross-period correlations, enabling a comprehensive representation of the coupled fuel-climate dynamics for FWI prediction.
Experimental results on real-world datasets demonstrate that FWI-MSNet achieves superior performance compared to existing methods, thereby providing a more reliable scientific foundation for forest fire prevention and management.
The terminology and abbreviations used in this paper are listed (Table 1).
FWI and its sub-indices FFMC, DMC, DC, ISI, and BUI are used to characterize fuel moisture and fire behavior potential, serving as core physical quantities for fuel‑climate coupling. 1D‑CNN and GRU are employed for multi‑scale feature extraction and temporal dependency modeling, respectively. These abbreviations are used throughout the model construction and feature selection process, providing a clear variable basis for physically aligned deep learning prediction.
2. Methodology
The FWI-MSNet framework consists of four main stages(Fig 2): data preparation, multi-scale feature construction, model construction, and model evaluation. First, multi-source datasets are cleaned, integrated, and normalized, followed by data augmentation. Second, 21 key features are organized into three groups according to the FWI system’s multi-scale temporal structure. Third, a parallel multi-scale 1D-CNN with kernel sizes of 7, 30, and 90 days extracts features at daily, monthly, and seasonal scales, which are then fused and processed by a GRU-Transformer collaborative mechanism. Fourth, the model is trained and evaluated through ablation experiments, comparisons with traditional models, and a case study on the 2020 Australian wildfires. The detailed methodology is elaborated in the following sections.
The workflow consists of four main stages: (1) Data Preparation, including data cleaning, FWI calculation, standardization, and augmentation; (2) Multi-scale Feature Construction, where 21 features are organized into three groups (short-term, medium-term, long-term) based on the FWI system’s temporal structure; (3) Model Construction, consisting of parallel multi-scale 1D-CNN, feature concatenation, GRU-Transformer collaboration, and fully connected output layer; and (4) Model Evaluation & Validation, including ablation experiments, comparisons with traditional models, and case study validation. Detailed parameter settings are provided in Tables 2–5 and Section 4.1.
2.1. Forest Fire Risk Weather Index
The FWI system is the physical core of this model, converting meteorological data into indices that reflect fuel moisture [24] and fire behavior [25]. Its chain includes three moisture and two behavior indices [26,27], leading to the final FWI [28]. These indices characterize drying conditions across time scales [29], providing a physical basis for multi-scale feature grouping. The mathematical expressions [30–32] are as follows:
Among them, is the humidity percentage of fine fuels;
is the rainfall correction term of DMC;
is the rainfall correction item of DC;
is the daily length coefficient;
is the daily length correction factor;
is the converted humidity index;
is the availability function of combustible materials;
is the temperature (°C);
is the relative humidity (%);
is the wind speed (km/h).
2.2. One-dimensional convolutional neural network
1D-CNN, a one-dimensional convolutional neural network [33], processes sequential data [34]. Its parallel configuration extracts multi-scale temporal features [35]. With different-sized convolution kernels [36], it captures local patterns from weather-driven responses to soil drought conditions [37], enabling hierarchical representation of physical grouping features. The mathematical expressions [38] are as follows:
Among them, represents the value of the output feature graph at position j;
is the input timing sequence;
is the weight of the convolutional kernel;
is the size of the convolutional kernel;
is the bias value;
is the index inside the convolutional kernel, from 0 to k − 1;
is the index of the output sequence.
2.3. Gated Recurrent Unit
GRU, a gated recurrent unit structure [39], is an efficient recurrent neural network variant [40]. Its update and reset gates alleviate gradient issues in long sequences [41] and model temporal dependencies [42]. In this model, the GRU layer models the evolution of multi-scale features, capturing mid-term memory and continuous patterns in environmental dynamics [43]. The mathematical expressions [44] are as follows:
In the formula: for the Sigmoid activation function, the output value is between [0, 1];
for the hyperbolic tangent activation function, the output value is between [−1, 1];
for the input vector of the current moment;
for the hidden state of the current moment;
for the output of the update gate;
for the loss of the reset gate Out;
、
and
are the weighted matrix;
、
and
are the bias term;
is the candidate hidden state.
2.4. Transformer
The Transformer’s self‑attention mechanism computes correlations between all sequence elements [45]. Placed after GRU, it extracts global contextual dependencies, enabling the model to capture both local temporal evolution and cross‑period fire‑risk connections for a comprehensive view of global risk patterns.
The core computational mechanisms [46] of the Transformer encoder include scaled dot-product attention, multi-head attention, and feed-forward networks [47], whose mathematical formulations and functions are described below. In addition, to compensate for the Transformer’s inherent inability to perceive sequence order, sinusoidal position encoding is added to the input embeddings.
Scaled dot-product attention is used to compute the correlation weights between queries and keys, and then perform a weighted sum of the values. The mathematical expressions [48] are as follows:
Multi-head attention projects the input into multiple subspaces and performs attention in parallel, enabling the model to jointly capture information from different representational subspaces. The mathematical expressions [48] are as follows:
where denote the query, key, and value matrices,
is the sequence length, and
is the dimension of each attention head.
Positional encoding is used to inject position information into the sequence to compensate for the Transformer’s inherent inability to perceive order. In this paper, sinusoidal functions are adopted to generate absolute positional encodings. The mathematical expressions [48] are as follows:
The feed-forward network applies a nonlinear transformation to the representation at each position, enhancing the model’s expressiveness. The mathematical expressions [48] are as follows:
2.5. Improved model
FWI‑MSNet (Fig 3) addresses multi‑scale fire‑risk capture limitations with a physically aligned framework that uses parallel 1D‑CNN channels for daily‑to‑quarterly feature extraction. GRU and Transformer [49] fuse features temporally and globally.
The input layer integrates multi-source fire risk data, including atmospheric humidity, soil data, and the FWI. Parallel multi-scale 1D-CNN channels with kernel sizes of 7, 30, and 90 extract short-term, medium-term, and long-term dynamic features at daily, monthly, and seasonal scales, respectively, to capture the multi-temporal-scale response of fuel–climate coupling. Subsequently, a GRU and a Transformer collaboratively fuse temporal evolution features with global contextual features to generate a fused feature map, and the prediction result is output through a fully connected layer. This framework provides a clear computational path for interpreting the multi-scale driving mechanisms of fire risk.
The proposed FWI-MSNet framework is implemented using Python 3.9 with PyTorch 1.10 as the deep learning backend. The implementation leverages the following libraries: NumPy for numerical computations, Pandas for data manipulation, scikit-learn for data preprocessing and evaluation metrics. The complete source code is available at the accompanying GitHub repository (https://github.com/CBJYB/FWIMSNet-Forest-Fire-Risk-Prediction-Framework). The overall execution flow of the model mainly includes the following steps.
- Data Preprocessing: standardization and data augmentation
- Construction of multi-scale input features
- Model Construction:
- Multi-scale Feature Extraction Module
- GRU-Transformer Collaborative Working Mechanism
- Output Layer Construction
- Model Training and Evaluation
- Model Prediction and Evaluation
2.5.1. Data preprocessing.
Normalization (Table 2) is performed using `sklearn.preprocessing.StandardScaler`, which ensures that each feature has zero mean and unit variance, thereby eliminating the influence of different scales. Data augmentation is implemented through sliding window sampling combined with noise injection. The base sequence length is set to 6 months with a sliding stride of 1, and multi-scale sequences of varying lengths are introduced to enrich sample diversity. Gaussian noise with standard deviations of 0.01 and 0.02 is applied to enhance the data, enabling better fitting of random fluctuations in the observed data.
The base sequence length is set to 6 months to cover a complete seasonal cycle of fuel moisture variation, consistent with the multi-scale temporal structure of the FWI system (Section 2.1). Multi-scale sequence lengths of 3, 4, 5, 7, and 8 months are introduced to enrich sample diversity by capturing both shorter and longer seasonal patterns. The stride is set to 1 to maximize sample diversity from the limited time series. Gaussian noise with standard deviations of 0.01 and 0.02 is added to improve model robustness against observation errors. The 80/20 chronological split ensures that training and test sets are strictly separated in time, preventing any future information from leaking into the training process.
2.5.2. Construction Of multi-scale input features.
This model groups features into three clear sets based on their physical processes and time scales [50]. It merges raw observations with calculated FWI indices [51] (Table 3) into a comprehensive feature pool that includes direct measurements like temperature and solar radiation, as well as physical simulation indices such as FFMC and DMC, using Pandas for data loading and scikit-learn’s StandardScaler for normalization.
The short-term (weekly) climate-driven group, represented by FFMC, captures the rapid response of fuel moisture to meteorological factors; the medium-term (monthly) eco-hydrological group, centered on DMC, reflects the seasonal moisture balance of the duff layer; the long-term (seasonal) soil-driven group, indicated by DC, characterizes the cumulative effect of deep-layer soil drought. This explicit grouping enables the model to extract features at time scales that match the fuel-climate coupling mechanism, providing a physical basis for the subsequent design of the multi-scale 1D-CNN kernel sizes (7, 30, and 90 days) and enhancing the model’s interpretability and generalization ability.
2.5.3. Multi-scale feature extraction module.
This module consists of three parallel one-dimensional convolutional layers (Table 4), each focusing on a specific time scale. One-dimensional convolution can effectively capture local dependencies by sliding along the temporal dimension, and its receptive field size is directly determined by the size of the convolution kernel. This module uses torch.nn to implement three parallel 1D convolutional layers, with Conv1d padding to maintain sequence length. Each convolution is followed by batch normalization (nn.BatchNorm1d) and ReLU activation (nn.ReLU), with dropout (nn.Dropout) for regularization.
The kernel sizes (7, 30, and 90 days) of the three parallel 1D‑CNN channels are explicitly mapped to three physical processes: weather-driven dynamics, eco-hydrological succession, and climate-scale drought accumulation, respectively. This design enables the model to extract features with receptive fields that match the fuel‑climate coupling mechanism: the short-term path captures the rapid response of fuel moisture to meteorological disturbances at the weekly scale, the medium-term path reflects the moisture balance of the duff layer at the monthly scale, and the long-term path characterizes the drought memory of deep-layer soil at the seasonal scale. All output feature maps have 32 dimensions, ensuring balanced weighting of features from different scales during subsequent fusion. This parameter configuration embeds domain priors into the network structure, reducing reliance on black-box learning.
The mathematical expressions [52] for the one-dimensional convolution operation on each path are as follows:
Among them, is the input feature sequence of the i-path;
and
is the trainable weight and bias of the path convolution kernel;
represents a one-dimensional convolution operation;
for the activation function, a nonlinear transformation is introduced;
∈RT*32 is the feature diagram of the output of the path.
The three parallel paths output feature maps, each with a dimension of 32. In order to integrate information from all scales, it is necessary to concatenate them on the feature dimension (the second dimension). The mathematical expression [53] is shown below.
represents the splicing operation. The fusion feature tensor
∈RT*96 also contains multi-scale information from weather transients, hydrological trends and climate background, providing a rich feature basis for subsequent time series modeling.
The feature extraction module (Fig 4) processes short‑, medium‑, and long‑term features through three parallel 1D‑CNN channels using kernel sizes 7, 30, and 90 to extract patterns corresponding to FFMC, DMC, and DC dynamics, respectively. The resulting feature maps are concatenated along the feature dimension to form a unified multi‑scale representation.
The kernel sizes 7, 30, and 90 correspond to the inherent variation periods (daily, monthly, and seasonal) of FFMC, DMC, and DC, respectively. This explicit mapping enables the network to extract features at time scales that match the physical mechanisms of fuel moisture, rather than blindly learning arbitrary temporal patterns. The feature maps output by the three parallel channels are concatenated to form a multi-scale representation, preserving cross-scale correlation information for the subsequent GRU-Transformer fusion.
2.5.4. GRU transformer collaborative working mechanism.
The GRU-Transformer collaborative mechanism [54] (Fig 5) uses a GRU layer to model temporal dependencies in multi-scale features, capturing mid-term evolution. The refined features are then processed by a Transformer encoder to mine global contextual correlations, and a fully connected layer generates the final fire risk prediction, enabling deep extraction from local to global insight. The mathematical expression is shown below.
The GRU layer first captures medium-term evolution patterns (e.g., continuous changes in fuel moisture), retaining sensitivity to short-term disturbances; then the Transformer encoder mines cross‑time‑step contextual dependencies from a global perspective (e.g., the remote modulation of climate anomalies on fire risk). This progressive “local→global” extraction enables the model to respond to both rapid fuel variations and slow climate trends within a physically consistent framework, avoiding the long‑term forgetting of a standalone RNN or the local over‑smoothing of a standalone Transformer.
where denotes the fused multi-scale features,
represents the GRU output capturing temporal dependencies,
is the output of the Transformer encoder, and
is the final FWI prediction,
denotes the fully connected layer.
The core hyperparameter configuration of the GRU‑Transformer collaborative module (Table 5) includes parameters such as input dimensions, the output of the multi‑scale CNN, GRU layer, Transformer encoder, and output layer.
The input dimension of 96 (32 dimensions from each of the three CNN channels) ensures lossless fusion of multi-scale features; both the GRU hidden layer and the Transformer d_model are set to 128 to facilitate collaboration; nhead = 8 and the feed-forward network dimension of 512 (4 times d_model) follow the standard Transformer configuration. The number of GRU layers is set to 1 to avoid overfitting in deep sequential networks; the number of Transformer layers is 2, which is sufficient to capture global context. The output layer reduces dimensions progressively (128, 64, 32, 1), compressing the fused features into a single-value FWI prediction. Overall, this configuration seamlessly connects the outputs of the physical grouping (Table 3) and multi-scale convolution (Table 4) to the temporal-global fusion module, without introducing an additional hyperparameter tuning burden.
3. Data processing and analysis
3.1. Data source and description
The Huitong Forest Ecological Station provides 18 years of continuous observations (2005–2022) across three core datasets (Table 6): climate [55], hydrological [56], and soil [57]. Climate data shows 96.51% completeness, soil data includes 82 variables with 66.04% completeness, and hydrological data contains 89 variables. These form the basis for fire risk and ecosystem coupling analysis.
Climate and environmental data have the highest completeness (96.51%), while soil and moisture data reach only 66.04% and 60.11%, respectively. This discrepancy stems from practical constraints in long‑term field observations: climate variables (temperature, humidity, etc.) are largely collected through automated systems, whereas soil and moisture data rely on manual or semi‑automatic measurements, resulting in more missing values. Within the fuel‑climate coupling framework, FWI and its sub‑indices (FFMC, DMC, DC) are highly dependent on the continuity of climate data; the high completeness of climate data provides reliable inputs for the model. In contrast, the lower completeness of soil and moisture features may introduce noise or bias. This suggests that appropriate imputation strategies should be adopted, and sensitivity analyses of low‑completeness features are necessary.
3.2. Data preprocessing process
Data cleaning and integration merged three datasets using year, month, day and ecological station code or plot code as keys into a unified table. Outliers were handled with the MAD method and ecological prior knowledge. To avoid data leakage, missing values were filled using a training‑set‑only K‑nearest neighbor algorithm: the dataset was first split chronologically into training (80%) and test (20%) sets, then the algorithm (k = 5, Euclidean distance) was fitted on the training set and applied to impute missing values in both sets. Feature engineering involved daily calculation of FWI system components FFMC, DMC, DC, ISI, BUI and FWI, as well as construction of temporal derivative features like consecutive dry days. From the initial 199 features, 21 were selected based on Pearson correlation with FWI (threshold |r| ≥ 0.6). No dimensionality reduction (e.g., PCA) was used, preserving physical interpretability. Data standardization performed Z-score normalization on all numerical features.
Where is the mean of the feature and
is the standard deviation.
3.3. Feature analysis and dataset construction
The three datasets contain 199 features. Correlation analysis (Fig 6) with FWI shows a concentrated distribution: 173 features have weak correlation between −0.1 and 0.1; three show very strong positive correlation between 0.8 and 1; 11 show very strong negative correlation between −0.8 and −1; and seven show strong negative correlation between −0.6 and −0.8. Negative correlations dominate, with monthly surface temperature showing the strongest correlation at −0.8550. Temperature-related indicators show strong negative correlation with FWI, highlighting the important role of temperature.
The vast majority of features (173 out of 199) show weak correlations, while strong correlations are concentrated among a few temperature-related indicators and are predominantly negative (the monthly land surface temperature correlation coefficient reaches −0.855). This phenomenon indicates that, within the fuel‑climate coupling framework, FWI is not a linear superposition of multiple factors but is strongly modulated by key climatic factors such as temperature. The dominance of weakly correlated features also reflects that a single linear correlation is insufficient to fully characterize fire risk, necessitating a deep learning model to capture nonlinear interactions. This provides a data-driven justification for the subsequent multi-scale, physically aligned architecture design.
This scatter plot shows the distribution of feature correlations, with many features densely distributed near zero, indicating weak correlations, while a few extremely strong correlations (Table 7) appear at both ends. This visualization confirms the central tendency in statistics and highlights rare but impactful key features, providing an intuitive basis for dataset construction.
Among the 21 strongly correlated features, positive correlations only appear for ISI, FFMC, and monthly mean air pressure; the remaining features show negative correlations, with radiation, soil temperature, DMC, DC, and other indicators exhibiting strong negative correlations with FWI (absolute values all > 0.73). This is consistent with the physical grouping in Table 3: ISI and FFMC belong to the short-term climate-driven group, reflecting the positive contribution of wind speed and surface fuel moisture to fire risk; whereas soil temperature (5 cm, 10 cm, 60 cm), DC, and others belong to the long-term soil-driven group, and their negative correlation suggests that soil heat accumulation is associated with lower fire risk. This correlation distribution provides a statistical basis for the physical mapping of the kernel sizes in Table 4 (7 days for short-term, 30 days for medium-term, and 90 days for long-term).
Due to the differences in the sizes of the three feature-grouped datasets, the three datasets were uniformly adjusted to 216 monthly records. To meet the model’s requirement for high temporal resolution, the monthly data were expanded to 6,570 daily records through linear interpolation. On this basis, a sliding window (window length of 180 days, step size of 1 day) was applied to generate 6,390 base sequence samples, and multi-scale sampling with window lengths of 90, 120, 150, 210, and 240 days was introduced, padded to 180 days. Gaussian noise augmentation with standard deviations of 0.01 and 0.02 was added, ultimately yielding 33,560 samples for model training and evaluation. For rigorous evaluation, all sequence samples generated in chronological order were directly split into the first 80% for training and the last 20% for testing.
4. Experiment and proof
To comprehensively evaluate the performance of the proposed FWI-MSNet model in forest fire risk prediction tasks, we designed and conducted a series of rigorous experiments.
4.1. Evaluation metrics
This study used four widely recognized indicators, RMSE, MAE, R2, and MAPE, to quantitatively evaluate the performance of the regression prediction task. The mathematical expressions are as follows:
Among them, is the total number of samples;
is the true value of the i sample;
is the predicted value of the i sample; and
is the evaluation value of the real value.
4.2. Model fitting analysis
The fitting process of the FWI-MSNet model is illustrated (Fig 7). The loss convergence curve (Fig 7a) shows that the initial loss of 10.3 decreases rapidly, enters a fluctuating convergence phase after about 20 rounds, and reaches −0.626 after 150 iterations. A contour map of parameter optimization (Fig 7b) shows weight parameters W1 and W2 migrating from a high‑loss area to a low‑loss area along the gradient descent path, with a final loss consistent with (Fig 7a). Together, they indicate fast and successful convergence and effective parameter updates via gradient descent in FWI-MSNet.
The loss curve (Fig 7a) shows that the initial loss of 10.3 decreases rapidly within 20 epochs, then enters a phase of slight fluctuations and convergence, indicating that the model can quickly capture the dominant gradient direction of the fuel‑climate coupling. After 150 epochs, the loss stabilizes at −0.626 with no obvious overfitting. In the parameter contour map (Fig 7b), W1 and W2 migrate smoothly from a high‑loss region to a global low‑loss region along the gradient descent path, consistent with the loss curve. This convergence behavior demonstrates that the collaborative structure of multi‑scale convolution and GRU‑Transformer does not increase optimization difficulty; instead, it provides a clear gradient descent path through physically aligned feature extraction.
4.3. Comparative experiments
To verify the effectiveness of each module in FWI-MSNet, this study designed two types of comparative experiments: ablation experiments and traditional model comparisons. All models predict the next 30 days. The ablation models include models without multi-scale convolution, without Transformer, without GRU, and without FWI features, all based on FWI-MSNet. The traditional models include XGBoost, LSTM, and CNN.
All models in this study were trained and evaluated under identical conditions to ensure fair comparison. Specifically, the same data split (chronological 80/20), feature normalization, batch size of 32, optimizer (Adam, initial learning rate 0.001), loss function (MSE), early stopping patience (20 epochs), and maximum epochs (200) were used for every deep learning model (including FWI-MSNet, its ablation variants, LSTM, and CNN). XGBoost was trained with grid-searched hyperparameters (number of trees = 100, max depth = 6, learning rate = 0.1).
The ablation experiment (Fig 8) shows model variants closely follow true FWI trends. Error metrics (Fig 8b) reveal significant differences: the Hybrid model achieves the lowest RMSE of 1.48 and MAE of 1.10. Most models (Fig 8c) show good fit with R2 values between 0.83 and 0.85, while the FWI Features model is lower at 0.512. MAPE comparisons (Fig 8d) range from 31.6% to 76.1%. Overall, these results demonstrate the superior performance of FWI-MSNet in FWI prediction.
FWI‑MSNet outperforms all variants in both RMSE (1.48) and MAE (1.10), while the model using only FWI features (without multi‑scale convolution and GRU‑Transformer) has an R2 as low as 0.512. This gap indicates that relying solely on the FWI index itself is insufficient to capture the nonlinear fuel‑climate coupling; multi‑scale feature extraction and temporal‑global fusion mechanisms are essential. Most variants achieve stable R2 values between 0.83 and 0.85, but the MAPE ranges widely (31.6%–76.1%), reflecting differences in prediction sensitivity to extreme fire risk values across architectures. Overall, these results validate the critical role of a physically aligned framework in improving prediction accuracy and robustness.
The traditional model comparison (Fig 9) shows that FWI-MSNet’s red prediction curve fits the true FWI values significantly better than CNN, XGBoost, and LSTM models (Fig 9a). FWI-MSNet achieves superior performance (Fig 9b) with an RMSE of 1.889, MAE of 1.405, R2 of 0.9251, and MAPE of 30.47%, far exceeding other models.
Compared with XGBoost, LSTM, and CNN, FWI‑MSNet achieves the highest agreement between its predicted curve and the true FWI values, with RMSE (1.889), MAE (1.405), and MAPE (30.47%) all significantly better than those of the other models, and an R2 of 0.9251. This advantage stems from the physically aligned design of the framework: although traditional models such as LSTM can capture temporal dependencies, they lack explicit multi‑scale convolution modeling of fuel moisture variations from daily to seasonal scales; XGBoost struggles with long‑range climate–fuel coupling. By leveraging parallel multi‑scale 1D‑CNN and GRU‑Transformer collaboration, FWI‑MSNet responds simultaneously to rapid fuel changes and slow climate trends, thus achieving superior performance in both extreme value fitting and trend tracking.
On the training set (Table 8), FWI-MSNet achieves the best performance (RMSE = 1.52, R2 = 0.94, MAPE = 28.2%). Removing any key component (multi-scale convolution, Transformer, GRU, or FWI features) significantly degrades all metrics, with the exclusion of FWI features causing the worst drop (R2 = 0.56, MAPE = 73.4%). Conventional models (XGBoost, LSTM, CNN) perform poorly, with XGBoost yielding a negative R2. These results confirm the effectiveness and necessity of each module in FWI-MSNet.
Removing multi‑scale convolution increases RMSE to 2.18, indicating that a single scale cannot simultaneously capture the dual response of fuel moisture to meteorological disturbances (weekly) and drought accumulation (seasonal). After removing the Transformer, R2 drops to 0.787, suggesting that global contextual dependencies (e.g., remote modulation by climate anomalies) are not effectively modeled. Removing GRU results in an MAE of 1.81, reflecting the loss of medium‑term temporal evolution information. The most severe performance degradation occurs when FWI features are removed (R2 = 0.556), confirming that the core fire risk index serves as a physical anchor for fuel‑climate coupling. Traditional models (XGBoost, LSTM, CNN) all achieve R2 values below 0.16 or even negative, due to their lack of physically aligned multi‑scale and collaborative mechanisms.
On the test set (Table 9), against seven models including no multi-scale convolution and XGBoost, FWI-MSNet reduced RMSE by 55.7%, MAE by 52.7%, and MAPE by 60.0%. Compared with each model, the arithmetic mean of RMSE, MAE, and MAPE improved by 56.1%, and the average absolute increase in R2 was 0.584, reflecting its significant superiority in prediction accuracy.
The comparative experimental results further validate the generalization advantage of FWI‑MSNet on the test set. Compared with the training set (Table 8), all models show increased errors, but FWI‑MSNet still maintains an RMSE of 1.69 and an R2 of 0.925, with the smallest degradation, indicating that the physically aligned design effectively suppresses overfitting. Removing FWI features causes R2 to drop to 0.512, the largest decline, highlighting the anchoring role of the core fire risk index in cross‑sample prediction. Removing the Transformer or the GRU increases RMSE to 3.00 and 2.45, respectively, demonstrating that temporal‑global collaboration is particularly critical for unknown fluctuations on the test set. Traditional models (XGBoost, CNN) still yield negative R2 values, suggesting that pure data‑driven methods lacking physical priors struggle to generalize.
Based on the evaluation results on the training and test sets, FWI-MSNet achieves an R2 of 0.9412 on the training set and 0.9251 on the test set, with RMSE only slightly increasing from 1.52 to 1.69 and MAPE marginally rising from 28.2% to 30.5%. No significant degradation is observed in any metric on the test set, and the performance trends of the ablation models and baseline models remain consistent. This indicates that FWI-MSNet has good fitting capability while effectively avoiding overfitting, demonstrating robust generalization performance.
The predicted curve (Fig 10) closely matches true FWI values and captures dynamic risk changes even during significant fluctuations, while prediction errors remain within ±2.5, demonstrating stable performance with practical value for forest fire prevention.
The prediction error remains within ±2.5 at all times. Even when the true FWI values exhibit significant fluctuations (e.g., days 15–25), the model responds promptly without lag. This indicates that the multi-scale features (daily to seasonal) extracted by the framework, together with the GRU-Transformer collaborative mechanism, effectively capture the rapid response and slow recovery of fuel moisture to climate anomalies. The stable error range also suggests that the model has the potential to provide reliable early warnings in real-world fire prevention scenarios.
4.4. Case study experiment
Australia’s catastrophic 2020 forest fires were triggered by lightning and human activities under extreme drought and windy conditions, burning 18.6 million hectares. To test FWI-MSNet’s robustness, this study used ERA5 data [58] and six representative areas covering diverse ecosystems. The FWI values for the case study are calculated using the same mathematical formulations described in Section 2.1, driven by ERA5 meteorological inputs (temperature, relative humidity, wind speed, and precipitation).
The FWI predictions for forest fires in Australia (Fig 11) show that the FWI-MSNet model effectively captures the evolution of fire risk—from low to extreme—across the six regions. During risk-level transitions, the predicted curve aligns with the actual trend, without significant lag. Quantitative evaluation further demonstrates the model’s robust cross-regional generalization: it achieves R2 values ranging from 0.57 to 0.77 and a stable MAPE of approximately 25% across all regions, despite being trained solely on data from Huitong, China. Notably, the model performs best in forest-dominated regions such as East Gippsland and the Blue Mountains, where R2 exceeds 0.76 (Table 10), while showing relatively lower performance in urban areas like Melbourne—a pattern consistent with the ecological similarity to the training domain. This case study confirms the model’s capability to accurately track fire risk dynamics with reliable transferability across regions, offering a valuable reference for forest fire prevention decision-making.
The model demonstrates cross-regional generalization capability in six regions of Australia. The model performs better in forest-dominated areas (e.g., East Gippsland, R2 > 0.76) than in urban areas (e.g., Melbourne), which is related to the regional consistency of the fuel‑climate coupling mechanism: forest areas have continuous fuel loads and moisture response patterns more similar to the training domain, whereas the surface conditions and anthropogenic disturbances in urban areas reduce the applicability of physical alignment. The MAPE remains stable at around 25%, and there is no significant lag in risk level transitions, indicating that the multi‑scale feature extraction has learned universal fire evolution patterns across ecological regions rather than over‑relying on local statistical characteristics.
Model performance in forest-dominated areas (East Gippsland, Blue Mountains) yields R2 values above 0.76, while in the urban area (Melbourne) it drops to 0.57. This discrepancy may be attributed to the forest cover characteristics of the training domain (Huitong Ecological Station, China): forest areas have high fuel continuity and moisture response patterns similar to Huitong, making the multi-scale coupling laws learned by the model more transferable; urban areas, however, are subject to disturbances from artificial underlying surfaces and fire suppression interventions, which weaken the effectiveness of physical alignment. The MAPE remains stable at around 25% across all regions, with RMSE ranging between 6.55 and 10.37, indicating that the model’s tracking of extreme fire danger level transitions is consistent, although attention should be paid to the impact of data scale (FWI value range) on error evaluation.
4.5. Insights from deep learning compared to alternative models
The novelty of this framework lies not in proposing a new deep learning architecture, but in establishing a physically aligned framework that explicitly couples the FWI system’s multi-scale structure with the model design. Beyond prediction accuracy, FWI-MSNet provides scientific insights that traditional models (XGBoost, LSTM, CNN) cannot offer. Based on the quantitative results reported in Table 9, three key findings emerge. First, the FWI system serves as a critical physical anchor for fire risk prediction: removing FWI features degrades R2 from 0.9251 to 0.5115, a 44.7% decrease relative to the full model. This confirms that the FWI system is not merely an additional input variable but a foundational physical prior that guides the learning process—a finding that traditional models, which treat all inputs uniformly, cannot reveal. Second, the simultaneous modeling of daily, monthly, and seasonal scales is essential: removing multi-scale convolution increases RMSE from 1.6892 to 2.4227, a 43.4% increase. This demonstrates that the fuel-climate coupling mechanism operates across multiple temporal scales simultaneously, and a single-scale model (e.g., standard CNN or LSTM) inevitably misses critical information at other scales. Third, the GRU and Transformer modules play distinct but complementary roles: removing GRU increases RMSE to 2.4466, while removing Transformer increases it to 3.0047, indicating that GRU primarily captures local sequential evolution of fuel moisture, whereas Transformer captures global cross-period correlations (e.g., remote modulation by climate anomalies). These insights, derived directly from the ablation experiments in Table 9, demonstrate that the physically aligned deep learning framework not only improves prediction accuracy but also provides interpretable understanding of the multi-scale fuel-climate coupling mechanism—a capability that conventional black-box models lack.
5. Discussion
Despite the promising results, this study has three main limitations. First, the model is trained exclusively on data from Huitong Ecological Station in China. Although its generalizability was tested on the 2020 Australian wildfire case study, validation across a broader range of geographic regions and climatic conditions remains necessary. Additionally, the current dataset lacks critical spatial information such as satellite-derived vegetation indices and topographic features (e.g., slope, aspect) that are known to influence fire behavior. Second, the current architecture operates on point-based time series without explicitly modeling spatial fire propagation across landscapes, and the GRU-Transformer mechanism introduces significant computational complexity. Third, although physically aligned in feature organization, the Transformer’s self-attention mechanism remains largely opaque, and the model does not output intermediate physical states (e.g., predicted moisture codes) that could be verified against physical equations, limiting full interpretability for domain experts.
Building on these limitations, future work will pursue three directions:
- (1). incorporating remote sensing data (Landsat, MODIS) and topographic variables to enable grid-based regional risk assessment;
- (2). exploring Graph Neural Networks (GNNs) to explicitly model fire spread across geographical grids while reducing computational costs via lightweight attention mechanisms;
- (3). developing a more interpretable model with auxiliary tasks to predict intermediate FWI components and physical loss functions that enforce known relationships (e.g., water balance constraints). Addressing these challenges will be crucial for evolving FWI-MSNet into a more robust, generalizable, and interpretable operational forecasting system.
6. Conclusions
This study proposed FWI-MSNet, a physically aligned deep learning framework for forest fire risk prediction based on 18 years of synchronous observation data from Huitong Ecological Station in China. The main findings are summarized as follows:
- (1). FWI-MSNet outperformed all seven baseline models, achieving an RMSE of 1.6892, MAE of 1.4046, R2 of 0.9251, and MAPE of 30.47%.
- (2). Physical alignment of features according to the FWI system’s multi-scale temporal structure substantially improved prediction accuracy; removing FWI features degraded R2 to 0.5115.
- (3). The model accurately captured the fire risk evolution trajectory in the 2020 Australian catastrophic wildfires, demonstrating potential generalizability beyond the training region.
References
- 1. Dong K, Gao Z, Fu L, Chen G, Ma X, Li N, et al. Research Progress on the Disaster Causing Mechanism of Forest Fires due to the Coupling of Multiple Factors. AJST. 2025;15(2):144–50.
- 2. Liu H, Shu L, Liu X, Cheng P, Wang M, Huang Y. Advancements in Artificial Intelligence Applications for Forest Fire Prediction. Forests. 2025;16(4):704.
- 3. Gao J, Wang L, Zhang W, Ning J, Li W, Hu T. Advances and Environmental Impact Assessment of Forest Fire Extinguishing Agents. Fire. 2025;8(11):411.
- 4. Wu G, Yao Q, Bai M, Shi L, Wang Z, Fang K. Weighing Policy Effectiveness Through Recent Forest Fire Status. Fire. 2024;7(12):432.
- 5.
Soliman H, Haque A. A wireless sensor network application in forest fire early detection: A smart and secure approach. In: 2024 Intelligent Systems and Machine Learning Conference (ISML), 2024. 106–11.
- 6. Khan T. Ultra-Low-Power Architecture for the Detection and Notification of Wildfires Using the Internet of Things. IoT. 2023;4(1):1–26.
- 7. Toledo-Castro J, Santos-González I, Caballero-Gil P, Hernández-Goya C, Rodríguez-Pérez N, Aguasca-Colomo R. Fuzzy-Based Forest Fire Prevention and Detection by Wireless Sensor Networks. Advances in Intelligent Systems and Computing. Springer International Publishing. 2018:478–88.
- 8. Moussa N, Nurellari E, Azbeg K, Boulouz A, Afdel K, Koutti L, et al. A reinforcement learning based routing protocol for software-defined networking enabled wireless sensor network forest fire detection. Future Generation Computer Systems. 2023;149:478–93.
- 9. Abbas K, Souane AA, Ahmad H, Suita F, Shu Z, Huang H, et al. Correlating Fire Incidents with Meteorological Variables in Dry Temperate Forest. Forests. 2025;16(1):122.
- 10. Chen J, Wang X, Yu Y, Yuan X, Quan X, Huang H. A novel fire danger rating model based on time fading precipitation model — A case study of Northeast China. Ecological Informatics. 2022;69:101660.
- 11.
Škurić Kuraži D, Nižetić Kosović I, Herceg Bulić I. Forest fire risk assessment with soil data in Croatia. Copernicus GmbH. 2022.
- 12. Abdelsattar M, Azim MA, AbdelMoety A, Emad-Eldeen A. Comparative analysis of deep learning architectures in solar power prediction. Sci Rep. 2025;15(1):31729. pmid:40877313
- 13. Marjani M, Mahdianpari M, Mohammadimanesh F. CNN-BiLSTM: A novel deep learning model for near-real-time daily wildfire spread prediction. Remote Sensing. 2024;16(8):1467.
- 14. Lin X, Li Z, Chen W, Sun X, Gao D. Forest fire prediction based on long- and short-term time-series network. Forests. 2023;14(4):778.
- 15. Miao X, Li J, Mu Y, He C, Ma Y, Chen J. Time Series Forest Fire Prediction Based on Improved Transformer. Forests. 2023;14(8):1596.
- 16. Santos LC, Lima MM, Bento VA, Nunes SA, DaCamara CC, Russo A, et al. An Evaluation of the Atmospheric Instability Effect on Wildfire Danger Using ERA5 over the Iberian Peninsula. Fire. 2023;6(3):120.
- 17. Kumari B, Pandey AC. Geo-informatics based multi-criteria decision analysis (MCDA) through analytic hierarchy process (AHP) for forest fire risk mapping in Palamau Tiger Reserve, Jharkhand state, India. J Earth Syst Sci. 2020;129(1).
- 18. Vidal-Silva C, Pizarro R, Castillo-Soto M, Ingram B, de la Fuente C, Duarte V, et al. A Comparative Study of a Deep Reinforcement Learning Solution and Alternative Deep Learning Models for Wildfire Prediction. Applied Sciences. 2025;15(7):3990.
- 19. Jiang W, Qiao Y, Su G, Li X, Meng Q, Wu H, et al. WFNet: A hierarchical convolutional neural network for wildfire spread prediction. Environmental Modelling & Software. 2023;170:105841.
- 20. Or D, Furtak-Cole E, Berli M, Shillito R, Ebrahimian H, Vahdat-Aboueshagh H, et al. Review of wildfire modeling considering effects on land surfaces. Earth-Science Reviews. 2023;245:104569.
- 21. Xu Z, Wang L, Cheng S, Rui X, Gao K, Zhu Y. Trustworthy Data-Driven Wildfire Risk Prediction and Understanding in Western Canada (Version 1). 2026.
- 22. Fu Y, Hu J, Song W, Cheng Y, Li R. Satellite observed response of fire dynamics to vegetation water content and weather conditions in Southeast Asia. ISPRS Journal of Photogrammetry and Remote Sensing. 2023;202:230–45.
- 23. Xu Z, Li J, Cheng S, Rui X, Zhao Y, He H, et al. Deep learning for wildfire risk prediction: Integrating remote sensing and environmental data. ISPRS Journal of Photogrammetry and Remote Sensing. 2025;227:632–77.
- 24. Xin X, Jiang H, Zhou G, Yu S, Wang Y. Canadian forest fire weather index (FWI) system: a review. Journal of Zhejiang A & F University. 2011.
- 25.
Škurić Kuraži D, Nižetić Kosović I, Herceg Bulić I. Forest fire risk assessment with soil data in Croatia. Copernicus GmbH. 2022.
- 26. Kussul N, Fedorov O, Yailymov B, Pidgorodetska L, Kolos L, Yailymova H, et al. Fire Danger Assessment Using Moderate-Spatial Resolution Satellite Data. Fire. 2023;6(2):72.
- 27. Matteo A, Garnés-Morales G, Moreno A, Ribeiro AFS, Azorin-Molina C, Bedia J, et al. Challenges in assessing Fire Weather changes in a warming climate. npj Clim Atmos Sci. 2025;8(1).
- 28.
Grillakis MG, Voulgarakis A, Rovithakis A, Seiradakis K, Koutroulis A, Field R. Ranking the sensitivity of climate variables and FWI sub-indices to global wildfire burned area. Copernicus GmbH. 2022.
- 29. Papagiannaki K, Giannaros TM, Lykoudis S, Kotroni V, Lagouvardos K. Weather-related thresholds for wildfire danger in a Mediterranean region: The case of Greece. Agricultural and Forest Meteorology. 2020;291:108076.
- 30.
Van Wagner CE, Pickett TL. Equations and FORTRAN program for the Canadian Forest Fire Weather Index System. Chalk River, Ontario: Canadian Forestry Service. 1985.
- 31. Wotton BM. Interpreting and using outputs from the Canadian Forest Fire Danger Rating System in research applications. Environ Ecol Stat. 2008;16(2):107–31.
- 32. Kudláčková L, Bartošová L, Linda R, Bláhová M, Poděbradská M, Fischer M, et al. Assessing fire danger classes and extreme thresholds of the Canadian Fire Weather Index across global environmental zones: a review. Environ Res Lett. 2024;20(1):013001.
- 33. Ahmadzadeh M, Zahrai SM, Bitaraf M. An integrated deep neural network model combining 1D CNN and LSTM for structural health monitoring utilizing multisensor time-series data. Structural Health Monitoring. 2024;24(1):447–65.
- 34. Alex SA, Jesu Vedha Nayahi J, Kaddoura S. Deep convolutional neural networks with genetic algorithm-based synthetic minority over-sampling technique for improved imbalanced data classification. Applied Soft Computing. 2024;156:111491.
- 35. Xu Y, Han L, Zhu T, Sun L, Du B, Lv W. Generic Dynamic Graph Convolutional Network for traffic flow forecasting. Information Fusion. 2023;100:101946.
- 36. Ali SW, Rashid MM, Yousuf MU, Shams S, Asif M, Rehan M. Towards the development of the clinical decision support system for the identification of respiration diseases via lung sound classification using 1D-CNN. Sensors. 2024;24(21):6887.
- 37. Liu L, Si Y-W. 1D convolutional neural networks for chart pattern classification in financial time series. J Supercomput. 2022;78(12):14191–214.
- 38. PyTorch Developers. torch.nn.Conv1d. PyTorch Documentation. https://pytorch.org/docs/stable/generated/torch.nn.Conv1d.html 2025. Accessed 2026 January 23.
- 39.
Kumar D, Aziz S. Performance Evaluation of Recurrent Neural Networks-LSTM and GRU for Automatic Speech Recognition. In: 2023 International Conference on Computer, Electronics & Electrical Engineering & their Applications (IC2E3), 2023. 1–6. https://doi.org/10.1109/ic2e357697.2023.10262561
- 40. Agarap AF. A neural network architecture combining gated recurrent unit (GRU) and support vector machine (SVM) for intrusion detection in network traffic data. arXiv. 2017.
- 41. Rana R. Gated Recurrent Unit (GRU) for Emotion Classification from Noisy Speech. arXiv. 2016.
- 42. Yang C-H, Molefyane T, Lin Y-D. The Forecasting of a Leading Country’s Government Expenditure Using a Recurrent Neural Network with a Gated Recurrent Unit. Mathematics. 2023;11(14):3085.
- 43. Li X, Ma X, Xiao F, Wang F, Zhang S. Application of Gated Recurrent Unit (GRU) Neural Network for Smart Batch Production Prediction. Energies. 2020;13(22):6121.
- 44.
Cho K, van Merrienboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, et al. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014. 1724–34. https://doi.org/10.3115/v1/d14-1179
- 45. Islam S, Elmekki H, Elsebai A, Bentahar J, Drawel N, Rjoub G, et al. A comprehensive survey on applications of transformers for deep learning tasks. Expert Systems with Applications. 2024;241:122666.
- 46. Vig J, Belinkov Y. Analyzing the Structure of Attention in a Transformer Language Model (Version 2). arXiv. 2019.
- 47. Wong M-F, Guo S, Hang CN, Ho SW, Tan CW. Natural language generation and understanding of big code for AI-assisted programming: A review. Entropy. 2023;25(6):888.
- 48. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention Is All You Need. arXiv. 2017:1. https://arxiv.org/abs/1706.03762
- 49. Zheng W, Zheng K, Gao L, Zhangzhong L, Lan R, Xu L, et al. GRU–Transformer: A Novel Hybrid Model for Predicting Soil Moisture Content in Root Zones. Agronomy. 2024;14(3):432.
- 50. Giannakopoulos C, Kostopoulou E, Varotsos KV, Tziotziou K, Plitharas A. An integrated assessment of climate change impacts for Greece in the near future. Reg Environ Change. 2011;11(4):829–43.
- 51.
Grillakis MG, Voulgarakis A, Rovithakis A, Seiradakis K, Koutroulis A, Field R. Ranking the sensitivity of climate variables and FWI sub-indices to global wildfire burned area. Copernicus GmbH. 2022.
- 52. Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. Commun ACM. 2017;60(6):84–90.
- 53. Du C, Wang Y, Wang C, Shi C, Xiao B. Selective feature connection mechanism: concatenating multi-layer CNN features with a feature selector. 2019. https://arxiv.org/abs/1811.06295
- 54. Mao W, Yu S, Chen W. Short-Term Power Load Forecasting Method Based on GRU-Transformer Combined Neural Network Model. CIT J comput inf technol. 2024;32(1):1–14.
- 55. National Research Station of Huitong Forest Ecosystems. A long-term monitoring dataset of meteorological indicators at National Research Station of Huitong Forest Ecosystems (2005-2022). National Ecosystem Science Data Center. 2025.
- 56. National Research Station of Huitong Forest Ecosystems. A long-term monitoring dataset of water environment indicators at National Research Station of Huitong Forest Ecosystems (2005-2022). National Ecosystem Science Data Center. 2025.
- 57. National Research Station of Huitong Forest Ecosystems. A long-term monitoring dataset of soil indicators at National Research Station of Huitong Forest Ecosystems (2005-2022). National Ecosystem Science Data Center. 2025.
- 58. Hersbach H, Bell B, Berrisford P. ERA5 hourly data on single levels from 1940 to present. Copernicus Climate Change Service (C3S) Climate Data Store. 2023.