Assessing crop yield and management yield zones using satellite imagery at the field scale
Files
TR Number
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Soybeans and corn are key crops in the United States, with many farms producing both. Predicting yield and delineating yield zones are crucial for nutrient management, guiding decisions to optimize production and increase yield. Accurate yield predictions are challenging due to variability in topography, soils, and management. We evaluated these factors using three (Edmunds, Hamlin, and Miner) South Dakota field trials and a set of six satellite images taken at different crop growth stages. Using Leave-One-Field-Out cross-validation, we trained an XGBoost machine learning model to predict yields and create spatial maps. Vegetation indices (VIs) from the R4/R5 stage showed the strongest correlation with soybean yield, especially Normalized Difference Vegetation Index (NDVI), Renormalized Difference Vegetation Index (RDVI), and Difference Vegetation Index (DVI), reaching 0.5-0.7 correlations. Soil Indices (SIs) had moderate correlation, with Coloration Index (CI) and Redness Index (RI) near 0.59-60 at Edmunds; Hamlin showed minimal soil-yield relationships. Topography showed a weak correlation, with elevation highest in 2019 (Pearson's R = 0.10) and 2021 (Pearson's R = 0.38). The best model at Edmunds (2021) had R2=0.54, RMSE=574 kg ha-1, and MAE=456 kg ha-1. Hamlin's performance was moderate; Miner performed poorly. Shapley Additive explanations (SHAP) analysis identified key predictors such as the DVI, RDVI, Green Chlorophyll Index (GCI), Triangular Greenness Index (TGI), and Green Normalized Difference Vegetation Index (GNDVI). In contrast, NDVI, Soil Adjusted Vegetation Index (SAVI), and Modified Soil Adjusted Vegetation Index 2 (MSAVI2) were deemed less significant. Combining topography, SIs, and VIs did not improve results. Focused on predicting high- and low-yielding zones using random forests (RF) and VIs, given that methods exist for delineating yield zones in crops. We hypothesized that integrating multi-year yield data, red-edge VIs, SIs, and topographic features (elevation, and slope) into a random forest (RF model) would improve soybean yield zone prediction. Predictors were tested across a range of Growing Degree Days (GDDs). Using Sentinel-2 satellite imagery from 2019 and 2021, twelve images were selected in a time series. An 80:20 split with k-5 cross-validation was used to train and test the RF model, evaluated by metrics like precision, F1-score, AUC-ROC, and accuracy. Error metrics were > 0.93, with best performance at GDD 1,080-1,300 and 1,300-1,500. Spatial maps were created to indicate yield potential expected under various weather conditions. To improve on the limitations of two-season data and to account for spatial-temporal variation, we collected historical yield data from Port Royal, VA, and delineated yield zones using clustering and spatial autocorrelation. We developed a workflow to clean, validate, and analyze multi-year data across crop rotations. Yield data from 78 site-years over 13 fields from 2018-2023 were filtered, aggregated into 10 m × 10 m grids, and spatially analyzed. Temporal stability facilitated classification of zones as high-stable, medium-stable, low-stable, or unstable, indicating long-term yield consistency. Unsupervised clustering enabled us to delineate management zones based on yield trends. We then compared average yields within zones and profitability potential, revealing distinct patterns with implications for site-specific management. Our results show that validated spatiotemporal yield maps effectively support zone classification and long-term decision-making for sustainable crop production.