Optical remote sensing (ORS) images serve as an essential Earth observation data source, capturing rich surface details that drive advancements in critical fields such as environmental monitoring, resource exploration, and disaster evaluation [1]. However, optical remote sensing images are often obscured by clouds. The International Satellite Cloud Climatology Project [2, 3] revealed that the average annual cloud cover worldwide is nearly 67%.
Early remote sensing data acquisition capabilities were once quite limited, prompting the initial development of single-image cloud removal techniques adapted from natural image dehazing that rely solely on the target optical image without auxiliary data, applying advanced processes such as spatial filtering, contrast enhancement, prior knowledge, and deep learning algorithms to excel in scenarios with light or evenly distributed cloud cover while performing poorly against thick or irregular clouds. To better capture and restore ground information hidden beneath thick clouds, multimodal cloud removal methods were developed by integrating complementary data sources, such as combining high-resolution multispectral data or fusing SAR data to leverage cloud-penetrating capabilities that improve the recovery of detailed texture features, making them effective for both thin and thick clouds and particularly adept at handling complex coverage scenarios despite being more intricate and potentially introducing additional noise.
Meanwhile, driven by advancements in remote sensing and image processing, multitemporal cloud removal techniques utilize multitemporal images as auxiliary data by fusing images of the same location captured at different intervals to distinguish dynamic cloud movements from stable ground features—using clear images from other time frames to assist in restoration—which excels in areas characterized by extensive or thick cloud coverage, although it faces challenges regarding positional alignment and remains unsuitable for regions with significant ground feature changes [4]. Figure 2 illustrates visual examples comparing the results of the single-image, multimodal, and multitemporal methods. In this article, we will focus specifically on single-image cloud removal approaches.
Single-Image Cloud Removal
Single-image cloud removal relies on information gathered from cloud-free areas to direct the restoration of details in obscured zones. By subduing cloud interference and highlighting surface characteristics, this approach effectively boosts overall image clarity. Furthermore, because it avoids the need to acquire or process extra data, it remains both cost-effective and time-efficient. Current techniques in this category are generally divided into statistical/physical model-based approaches and deep learning-based strategies [5].
Statistical and physical model-based methods
Statistical and physical model-based methods operate on the principle that distinctive statistical patterns or physical property differences exist between valid and missing data within a cloudy image, allowing these discrepancies to be leveraged for cloud separation and image restoration. By utilizing physical imaging models alongside spatial correlations and frequency variances between affected and clear areas, these traditional approaches effectively reconstruct obscured sections.
Spatial-based methods [6], [7] rely on the assumption that valid and missing data share comparable textural details or statistical traits, using this inherent similarity to estimate missing content and generate a visually coherent cloud-free result. However, these techniques are typically restricted to handling only speckled clouds, and the semantic details recovered beneath heavy cloud layers frequently drift far from the true original data.
Frequency-based methods [8], [9], [10] focus on the low-frequency properties characteristic of thin clouds, applying designed low-pass filters to isolate and subsequently suppress the cloud layer. Even so, finding the exact optimal cutoff frequency remains a notoriously difficult task, often resulting in the accidental loss of essential low-frequency data in otherwise clear regions.
Physical-based methods closely follow atmospheric imaging principles, carefully accounting for phenomena like atmospheric absorption and scattering to clear away thin clouds through precise parameter adjustment [4]. Prior-based frameworks, for example, estimate the global atmospheric light A and transmission map t(x) under specific assumptions to retrieve the underlying clear image J(x) using the atmospheric scattering model defined by I(x) = J(x)t(x) + A(1 - t(x)) [11].
While foundational techniques like the uniform atmospheric correction by Zhang et al. [12] and the popular dark channel prior (DCP) introduced by He et al. [13] prove highly effective for thin cloud removal, any inaccuracy in estimating these atmospheric parameters can severely undermine the results, leaving behind residual clouds or unwanted artifacts [14].
Ultimately, while statistical and physical model-based methods offer strong theoretical clarity and perform reliably when dealing with uniform thin clouds—relying on multi-dimensional pattern recognition or parameter estimation—they often fall short under complex, shifting meteorological conditions where strict underlying assumptions break down [4].
Deep Learning-Based Methods
Propelled by rapid advancements in deep learning, single-image cloud removal has unlocked powerful new capabilities. By constructing deep neural networks to master the intricate mappings between cloudy inputs and clear targets, these techniques eliminate cloud blockages and boost overall image quality with unprecedented accuracy, representing a major milestone in single optical remote sensing (ORS) image restoration.
Convolutional Neural Networks (CNNs): Powered by robust data-fitting strengths, CNNs excel at wiping out thin clouds by learning nonlinear transformations. To maximize visual fidelity, specialized architectures incorporate residual blocks to combat additive noise, multiscale convolutions to capture fine-grained textures, and attention mechanisms to focus on crucial features. Even so, the constrained kernel sizes of standard CNNs limit global-local context interaction, occasionally leaving behind artifacts or cloud residues.
U-Shaped Encoder-Decoder Architectures: Many frameworks employ U-Net structures to rebuild clear landscapes by linking encoding steps directly to decoding counterparts via skip connections. Because direct end-to-end mapping lacks transparency, some models fuse CNNs with frequency-domain properties. Nonetheless, traditional U-Net models often fail to harmonize low-level and high-level traits smoothly, leading to overlooked details, while their reliance on massive sets of synthetic paired training data can weaken real-world generalization.
Generative Adversarial Networks (GANs): Serving as powerful unsupervised tools to model complex data distributions, GANs pair a generator with a discriminator to master cloudy patterns from unlabeled sources, reducing reliance on heavy annotation. While unidirectional GANs offer strong real-world flexibility, their performance remains bottlenecked by the absence of clean ground-truth labels. To counter this, CycleGAN architectures [15], [16], [17] establish two-way cross-domain mappings managed by cycle-consistency and identity losses to safeguard color and texture integrity. However, CycleGAN struggles to capture global data distributions thoroughly, causing asymmetric domain translation and less-than-ideal texture recovery beneath heavy clouds.
Ultimately, while deep learning-driven cloud removal boasts exceptional feature extraction capabilities, it is still constrained by heavy data dependencies, high computational demands, and an inherent inability to recover information entirely swallowed by thick cloud cover.
Future Work
Single-image cloud removal techniques offer distinct advantages thanks to their low reliance on extra auxiliary data, which greatly cuts down the costs and complexities tied to data acquisition. While they perform well when dealing with thin clouds, these approaches struggle to clear away dense clouds and shadows, frequently causing artifacts or color distortion. Furthermore, their accuracy and robustness often drop when faced with complex cloud scenes.
Looking ahead, as computer vision and remote sensing continue to evolve, single-image cloud removal is primed for major breakthroughs:
- Advanced Neural Network Models: Developing more sophisticated architectures will allow models to better distinguish between ground objects and clouds, yielding higher accuracy.
- GAN-Based Unpaired Learning: Leveraging these mechanisms will help boost the overall generalization capabilities of the models.
- Multitask Integration: Building a more holistic processing pipeline through the integration of multitask learning with other cutting-edge computer vision technologies will further enhance ORS image processing.
References
- [1] M. Li, Q. Xu, J. Guo, and W. Li, “DecloudNet: Cross-patch consistency is a nontrivial problem for thin cloud removal from wide-swath multispectral images,” IEEE Trans. Geosci. Remote Sens., vol. 62, 2024, Art. no. 5407614.
- [2] Zhang, Y.; Rossow, W.B.; Lacis, A.A.; Oinas, V.; Mishchenko, M.I. Calculation of radiative fluxes from the surface to top of atmosphere based on ISCCP and other global data sets: Refinements of the radiative transfer model and the input data. J. Geophys. Res. Atmos. 2004, 109, D19.
- [3] King, M.D.; Platnick, S.; Menzel, W.P.; Ackerman, S.A.; Hubanks, P.A. Spatial and temporal distribution of clouds observed by MODIS onboard the Terra and Aqua satellites. IEEE Trans. Geosci. Remote Sens. 2013, 51, 3826–3852.
- [4] Ning, Jin & Xie, Lianbin & Yin, Jie & Liu, Yiguang. (2025). Cloud Removal Advances: A Comprehensive Review and Analysis for Optical Remote Sensing Images. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing. 18. 15914-15930. 10.1109/JSTARS.2025.3580718.
- [5] W. Li, Y. Li, D. Chen, and J. C.-W. Chan, “Thin cloud removal with residual symmetrical concatenation network,” ISPRS J. Photogrammetry Remote Sens., vol. 153, pp. 137–150, 2019.
- [6] M. Xu, X. Jia, M. Pickering, and S. Jia, “Thin cloud removal from optical remote sensing images using the noise-adjusted principal components transform,” ISPRS J. Photogrammetry Remote Sens., vol. 149, pp. 215–225, 2019.
- [7] C. Yu, L. Chen, L. Su, M. Fan, and S. Li, “Kriging interpolation method and its application in retrieval of modis aerosol optical depth,” in Proc. 19th Int. Conf. Geoinformatics, IEEE, 2011, pp. 1–6.
- [8] G. Hu, X. Li, and D. Liang, “Thin cloud removal from remote sensing images using multidirectional dual tree complex wavelet transform and transfer least square support vector regression,” J. Appl. Remote Sens., vol. 9, no. 1, 2015, Art. no. 095053.
- [9] M. Wan and X. Li, “Removing thin cloud on single remote sensing image based on SWF,” in Proc. IEEE Int. Conf. Online Anal. Comput. Sci. (ICOACS), IEEE, 2016, pp. 397–400.
- [10] S. Zhao and S. Dai, “Automated cloud removal and filling in optical remote sensing images,” in Proc. Int. Conf. Virtual Reality Visual. (ICVRV), IEEE, 2016, pp. 292–297.
- [11] J. Ning, Y. Zhou, X. Liao, and B. Duo, “Single remote sensing image dehazing using robust light-dark prior,” Remote Sens., vol. 15, no. 4, 2023, Art. no. 938.
- [12] Y. Zhang, B. Guindon, and J. Cihlar, “An image transform to characterize and compensate for spatial variations in thin cloud contamination of landsat images,” Remote Sens. Environ., vol. 82, no. 2/3, pp. 173–187, 2002.
- [13] K. He, J. Sun, and X. Tang, “Single image haze removal using dark channel prior,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 33, no. 12, pp. 2341–2353, Dec. 2011.
- [14] S. Shi, Y. Zhang, X. Zhou, and J. Cheng, “A novel thin cloud removal method based on multiscale dark channel prior (MDCP),” IEEE Geosci. Remote Sens. Lett., vol. 19, 2021, Art. no. 1001905.
- [15] P. Singh and N. Komodakis, “Cloud-GAN: Cloud removal for sentinel2 imagery using a cyclic consistent generative adversarial networks,” in Proc. IEEE Int. Geosci. Remote Sens. Symp., IEEE, 2018, pp. 1772–1775.
- [16] Y. Zi, F. Xie, X. Song, Z. Jiang, and H. Zhang, “Thin cloud removal for remote sensing images using a physical-model-based CycleGAN with unpaired data,” IEEE Geosci. Remote Sens. Lett., vol. 19, 2021, Art. no. 1004605.
- [17] Y. Mo, C. Li, Y. Zheng, and X. Wu, “DCA-CycleGAN: Unsupervised single image dehazing using dark channel attention optimized CycleGAN,” J. Vis. Commun. Image Representation, vol. 82, 2022, Art. no. 103431.