A Standardized Framework for Cleaning Non-Normal Yield Data from Wheat and Barley Crops, and Validation Using Machine Learning Models for Satellite Imagery
Rattachement africain : es. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Modern combine harvesters can collect real-time geolocated yield data, but it is subject to errors. Various protocols have been proposed to clean this data, each with varying levels of complexity. This data is valuable for precision agriculture to implement site-specific management and to train models to predict yield using remote sensing data. Machine learning and deep learning techniques have shown their potential for precision agriculture, and their performance shows no significant differences between models trained with data cleaned using a computationally demanding protocol or a simpler one, such as parametric filtering. However, parametric filtering approaches primarily rely on statistics that are highly sensitive to data distribution and do not effectively filter inliers. The objective of this study is to develop a data-cleansing method that leverages robust statistical measures, specifically the median and interquartile range, to effectively identify and filter outliers and inliers while retaining valid observations in datasets collected from combine harvesters, thereby minimizing the influence of non-normal data distributions. Different levels of data cleaning were applied to a total of 7399 ha of wheat and barley crops, and the quality of each cleaning level was compared. The selected protocol improved the spatial structure of the data, deleting up to 42% and 33% of the data at the polygon level, for wheat and barley, respectively. It increased the mean and median, and decreased the standard deviation and coefficient of variation of the data. Between 78.7% and 82.9% of the fields showed a normal distribution after applying the selected method, and machine learning performance improved compared with the raw data. Compared with previous data cleaning studies, the present work proposes an automatic, low-computational, parametric filtering method that uses robust statistics for non-normal distributions. In addition, its scalability has been demonstrated by applying the method to a large dataset, improving data quality and the performance of yield-prediction ML models in all cases.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Standardized Framework for Cleaning Non-Normal Yield Data from Wheat and Barley Crops, and Validation Using Machine Learning Models for Satellite Imagery
- Date Crossref
- 05/02/2026
- Éditeur
- MDPI AG
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Universitat Politècnica de València pays non établi dans la noticeUniversité ou école supérieure
-
Centro de Investigación del Regadío y Agrosistemas Mediterráneos pays non établi dans la noticeInstitution
-
Departamento de Física Aplicada pays non établi dans la noticeInstitution
-
Centro de Tecnologías Físicas pays non établi dans la noticeInstitution
-
Departamento de Matemática Aplicada pays non établi dans la noticeInstitution
-
Instituto Universitario de Investigación de Matemática Multidisciplinar pays non établi dans la noticeStructure de recherche
-
Departamento de Producción Vegetal pays non établi dans la noticeInstitution
Universitat Politècnica de València, Centro de Investigación del Regadío y Agrosistemas Mediterráneos et Departamento de Física Aplicada, avec 4 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.