Comparative Evaluation of Machine Learning Models for Satellite Chlorophyll-a Gap Reconstruction in the Chesapeake Bay
Rakshita Chidananda, Anusha Srirenganathan Malarvizhi, Samir Ahmed, Elena Zhang, Chaowei Phil Yang
Harmful algal blooms (HABs) are increasing in frequency in the Chesapeake Bay, posing risks to marine ecosystems, water quality, and public health. Chlorophyll-a (Chl-a) is a widely used indicator of algal biomass, and satellite observations such as Sentinel-3 Ocean and Land Color Instrument (OLCI) enable large-scale monitoring of bloom dynamics. However, cloud cover and atmospheric interference frequently introduce missing pixels in daily satellite products, reducing temporal continuity and limiting monitoring reliability. Satellite-derived chlorophyll-a (Chl-a) data exhibit substantial missingness, with daily pixel gaps ranging from approximately 52.30% to 100% (mean ≈ 88.95%). This study evaluates spatial interpolation, EOF-based, supervised machine-learning, deep-learning, and convolutional autoencoder approaches for reconstructing missing Chl-a values. Sentinel-3 OLCI Chl-a data from 2023–2024 were used for model training, while data from 2025 served as a temporally independent test set to avoid spatiotemporal leakage. To simulate cloud-induced data gaps, artificial missingness scenarios ranging from 50% to 90% were applied for the Inverse Distance Weighting (IDW) and Data Interpolating Empirical Orthogonal Functions (DINEOF) baseline approaches, while machine-learning, deep-learning, and convolutional autoencoder models were evaluated using real satellite-derived missing observations. The evaluated models include IDW, DINEOF, K-Nearest Neighbors (KNN), Random Forest (RF), Extra Trees (ET), XGBoost, a Long Short-Term Memory (LSTM) network, and a Temporal Data Interpolating Convolutional Autoencoder (Temporal DINCAE). Model performance was assessed using Root Mean Square Error (RMSE), Mean Absolute Error (MAE), prediction bias, and the coefficient of determination (R2). Results indicate that tree-based ensemble models outperform spatial interpolation and EOF-based methods, with XGBoost achieving the best overall performance (R2 ≈ 0.86; RMSE ≈ 9.61 mg m−3). The LSTM model achieved lower prediction errors (RMSE ≈ 5.87 mg m−3; MAE ≈ 2.16 mg m−3), highlighting the benefit of incorporating temporal dependencies, although with slightly reduced variance capture. The convolutional autoencoder-based Temporal DINCAE model achieved strong reconstruction performance (R2 ≈ 0.84; RMSE ≈ 11.15 mg m−3). Uncertainty quantification shows that Extra Trees tends to underestimate uncertainty with narrower prediction intervals, whereas XGBoost provides better-calibrated but wider intervals.