CORTEXA
← Browse
crossrefAgronomy2026-02-05Cited by 1

A Standardized Framework for Cleaning Non-Normal Yield Data from Wheat and Barley Crops, and Validation Using Machine Learning Models for Satellite Imagery

Patricia Arizo-García, Sergio Castiñeira-Ibáñez, Enric Cruzado-Campos, Beatriz Ricarte, Constanza Rubio, Alberto San Bautista

Modern combine harvesters can collect real-time geolocated yield data, but it is subject to errors. Various protocols have been proposed to clean this data, each with varying levels of complexity. This data is valuable for precision agriculture to implement site-specific management and to train models to predict yield using remote sensing data. Machine learning and deep learning techniques have shown their potential for precision agriculture, and their performance shows no significant differences between models trained with data cleaned using a computationally demanding protocol or a simpler one, such as parametric filtering. However, parametric filtering approaches primarily rely on statistics that are highly sensitive to data distribution and do not effectively filter inliers. The objective of this study is to develop a data-cleansing method that leverages robust statistical measures, specifically the median and interquartile range, to effectively identify and filter outliers and inliers while retaining valid observations in datasets collected from combine harvesters, thereby minimizing the influence of non-normal data distributions. Different levels of data cleaning were applied to a total of 7399 ha of wheat and barley crops, and the quality of each cleaning level was compared. The selected protocol improved the spatial structure of the data, deleting up to 42% and 33% of the data at the polygon level, for wheat and barley, respectively. It increased the mean and median, and decreased the standard deviation and coefficient of variation of the data. Between 78.7% and 82.9% of the fields showed a normal distribution after applying the selected method, and machine learning performance improved compared with the raw data. Compared with previous data cleaning studies, the present work proposes an automatic, low-computational, parametric filtering method that uses robust statistics for non-normal distributions. In addition, its scalability has been demonstrated by applying the method to a large dataset, improving data quality and the performance of yield-prediction ML models in all cases.

View free PDFSource page

Related papers

crossrefAgronomy2023-12-19Cited by 32

Prediction of Stem Water Potential in Olive Orchards Using High-Resolution Planet Satellite Images and Machine Learning Techniques

Simone Pietro Garofalo, Vincenzo Giannico, Leonardo Costanza, Salem Alhajj Ali, Salvatore Camposeo, Giuseppe Lopriore, et al.

Assessing plant water status accurately in both time and space is crucial for maintaining satisfactory crop yield and quality standards, especially in the face of a changing climate. Remote sensing technology offers a promising alternative to traditional in situ measurements for…

View free PDFSource page
crossrefAgronomy2024-12-17Cited by 58

Machine Learning and Deep Learning for Crop Disease Diagnosis: Performance Analysis and Review

Habiba Njeri Ngugi, Andronicus A. Akinyelu, Absalom E. Ezugwu

Crop diseases pose a significant threat to global food security, with both economic and environmental consequences. Early and accurate detection is essential for timely intervention and sustainable farming. This paper presents a review of machine learning (ML) and deep learning (…

View free PDFSource page
crossrefAgronomy2025-10-31

Quantifying Grazing Intensity from Aboveground Biomass Differences Using Satellite Data and Machine Learning

Ritu Su, Yong Yang, Shujuan Chang, Gudamu A, Xiangjun Yun, Xiangyang Song, et al.

Accurately quantifying grazing intensity (GI) is crucial for assessing grassland utilization and supporting sustainable management. Traditional livestock-based approaches cannot capture the spatial heterogeneity of grazing or its dynamic response to climate variability. The objec…

View free PDFSource page
crossrefAgronomy2025-03-20Cited by 10

Breeding of Solanaceous Crops Using AI: Machine Learning and Deep Learning Approaches—A Critical Review

Maria Gerakari, Anastasios Katsileros, Konstantina Kleftogianni, Eleni Tani, Penelope J. Bebeli, Vasileios Papasotiropoulos

This review discusses the potential of artificial intelligence (AI), particularly machine learning (ML) and its subset, deep learning (DL), in advancing the genetic improvement of Solanaceous crops. AI has emerged as a powerful solution to overcome the limitations of traditional…

View free PDFSource page
crossrefAgronomy2024-03-08Cited by 17

An Estimation of the Leaf Nitrogen Content of Apple Tree Canopies Based on Multispectral Unmanned Aerial Vehicle Imagery and Machine Learning Methods

Xin Zhao, Zeyi Zhao, Fengnian Zhao, Jiangfan Liu, Zhaoyang Li, Xingpeng Wang, et al.

Accurate nitrogen fertilizer management determines the yield and quality of fruit trees, but there is a lack of multispectral UAV-based nitrogen fertilizer monitoring technology for orchards. Therefore, in this study, a field experiment was conducted by UAV to acquire multispectr…

View free PDFSource page
crossrefAgronomy2026-03-14

Regional-Scale Mapping of Gully Network in Mediterranean Olive Landscapes Using Machine Learning Algorithms: The Guadalquivir Basin

Paula González-Garrido, Adolfo Peña-Acevedo, Francisco-Javier Mesas-Carrascosa, Juan Julca-Torres

Gully erosion is a significant threat to the sustainability of soil in Mediterranean basins. Despite its impact, there is a lack of research providing accurate regional-scale cartography of complete gully networks. This study aims to automatically map the gully network in the oli…

View free PDFSource page