Filling Data Gaps: Comparing Regression-based Models and Machine Learning Methods for Rainfall Time Series Reconstruction
Rattachement africain : br. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract Rainfall time series are essential for hydrological and climate studies; however, data scarcity remains a persistent challenge that compromises the reliability of analyses and modeling. This study compares the performance of regression based models and machine learning (ML) methods for gap filling in monthly rainfall data from Northern Minas Gerais, Brazil, a region characterized by high climatic variability and limited monitoring infrastructure. Ten missing data levels, ranging from 5% to 50%, were simulated using historical records from seven rainfall stations. Four regression based approaches, Regional Weighting (RW), Simple Linear Regression (SLR), Multiple Linear Regression (MLR), and Regression Based Weighting (RWR), were evaluated alongside three ML methods, Random Forest (RF), Support Vector Machines (SVM), and k Nearest Neighbors (KNN). Model performance was assessed using Root Mean Square Error (RMSE), Symmetric Mean Absolute Percentage Error (SMAPE), Nash Sutcliffe Efficiency (NSE), and the coefficient of determination (R 2 ), with nonparametric statistical tests applied to identify significant differences. The results showed that regression and weighting based models consistently outperformed ML methods across all missing data levels. RWR and RW achieved the best overall performance, with the lowest RMSE and SMAPE values and the highest NSE and R 2 values, and no statistically significant differences were observed between them. MLR performed well only at moderate levels of missing data, with significant performance degradation beyond 40% of missing values. Among ML methods, RF and SVM showed intermediate performance, while KNN and SLR yielded the weakest results. These findings highlight that weighting schemes based on the relevance and proximity of neighboring stations provide more robust rainfall estimates than complex nonlinear models in data scarce and highly variable climatic regions. Graphical Abstract The graphical abstract presents the workflow used to evaluate the performance of gap-filling techniques in monthly rainfall time series from Northern Minas Gerais, Brazil. The process begins with historical rainfall data, in which artificial gaps ranging from 5% to 50% are systematically introduced to simulate missing values. These incomplete time series are then processed using seven gap filling models that include both statistical and machine learning approaches, namely Regional Weighting (RW), Simple Linear Regression (SLR), Multiple Linear Regression (MLR), Regression Based Weighting (RWR), Random Forest (RF), Support Vector Machines (SVM), and K Nearest Neighbors (KNN). Each technique is applied iteratively across 100 simulations to ensure statistical reliability. The performance of the reconstructed series is evaluated using four metrics, Root Mean Square Error (RMSE), Symmetric Mean Absolute Percentage Error (SMAPE), Nash Sutcliffe Efficiency (NSE), and the coefficient of determination (R 2 ). The flowchart visually synthesizes this stepwise approach, covering data preprocessing, gap simulation, model application, and performance comparison, and highlights the systematic strategy adopted in the study to identify the most accurate gap filling models under different levels of missing data.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Filling Data Gaps: Comparing Regression-based Models and Machine Learning Methods for Rainfall Time Series Reconstruction
- Date Crossref
- 21/07/2026
- Éditeur
- Springer Science and Business Media LLC
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.