Cross-Modal Generative Models for Remote Sensing Image Enhancement and Earth Observation Applications
Keywords:
cross-modal generative models, remote sensing, Earth observation, image enhancement, deployment architectures, data governance, foundation models, fairness, sustainabilityAbstract
Cross-modal generative models are becoming a central architectural paradigm for remote sensing image enhancement and large-scale Earth observation workflows. These models integrate heterogeneous data sources, including multispectral imagery, synthetic aperture radar, text-based environmental reports, and geospatial metadata, to reconstruct degraded scenes, fill missing observations, and generate physically plausible interpretations of the Earth surface. This paper provides a systems-level analysis of cross-modal generative approaches in remote sensing, focusing not on narrow algorithmic performance but on structural trade-offs, data governance, deployment architectures, robustness, fairness, sustainability, and policy implications. It examines how adversarial and diffusion-based generative mechanisms are reorganized within Earth observation pipelines and how they interact with transformer-based representation learning. The discussion emphasizes the importance of multi-sensor data governance, calibration under distribution shift, edge-to-cloud deployment, and the regulatory pressures emerging around foundation models. The paper further considers how generative outputs are increasingly used in humanitarian mapping, climate monitoring, and environmental policy, requiring careful attention to uncertainty communication and equitable access. By connecting architectural decisions to broader infrastructural and societal constraints, the paper argues that the next generation of Earth observation systems must treat generative enhancement not merely as an image reconstruction task but as a governed socio-technical infrastructure.
References
1. Zhu, X. X., Tuia, D., Mou, L., Xia, G.-S., Zhang, L., Xu, F., & Fraundorfer, F. (2017). Deep learning in remote sensing: A comprehensive review and list of resources. IEEE Geoscience and Remote Sensing Magazine, 5(4), 8–36.
2. Ma, L., Liu, Y., Zhang, X., Ye, Y., Yin, G., & Johnson, B. A. (2019). Deep learning in remote sensing applications: A meta-analysis and review. ISPRS Journal of Photogrammetry and Remote Sensing, 152, 166–177.
3. Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27, 2672–2680.
4. Isola, P., Zhu, J.-Y., Zhou, T., & Efros, A. A. (2017). Image-to-image translation with conditional adversarial networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1125–1134.
5. Zhu, J.-Y., Park, T., Isola, P., & Efros, A. A. (2017). Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference on Computer Vision, 2223–2232.
6. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10684–10695.
7. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.
8. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
9. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations.
10. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning, 8748–8763.
11. Cong, Y., Khanna, S., Meng, C., Liu, P., Rozi, E., He, Y., Burke, M., Lobell, D., & Ermon, S. (2022). SatMAE: Pre-training transformers for temporal and multi-spectral satellite imagery. Advances in Neural Information Processing Systems, 35, 197–211.
12. Sun, X., Wang, P., Lu, W., Zhu, Z., Lu, X., He, Q., Li, J., Rong, X., Yang, Z., Chang, H., & Fu, K. (2022). RingMo: A remote sensing foundation model with masked image modeling. IEEE Transactions on Geoscience and Remote Sensing, 60, 1–22.
13. Bastani, F., Wolters, P., Gupta, R., Ferdinando, J., & Kembhavi, A. (2023). SatlasPretrain: A large-scale dataset for remote sensing image understanding. Proceedings of the IEEE/CVF International Conference on Computer Vision, 16772–16782.
14. Jean, N., Burke, M., Xie, M., Davis, W. M., Lobell, D. B., & Ermon, S. (2016). Combining satellite imagery and machine learning to predict poverty. Science, 353(6301), 790–794.
15. Reichstein, M., Camps-Valls, G., Stevens, B., Jung, M., Denzler, J., Carvalhais, N., & Prabhat. (2019). Deep learning and process understanding for data-driven Earth system science. Nature, 566(7743), 195–204.
16. Tuia, D., Roscher, R., Wegner, J. D., Jacobs, N., Zhu, X. X., & Camps-Valls, G. (2021). Toward a collective agenda on AI for Earth science data analysis. IEEE Geoscience and Remote Sensing Magazine, 9(2), 88–104.
17. Chen, Ce, et al. "JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators." arXiv preprint arXiv:2606.28421 (2026).
18. Lacoste, A., Luccioni, A., Schmidt, V., & Dandres, T. (2019). Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700.
19. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650.
20. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.
21. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
22. Veale, M., & Binns, R. (2017). Fairer machine learning in the real world: Mitigating discrimination without collecting disparate impact. Big Data & Society, 4(2), 2053951717743530.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Engineering Systems and Digital Innovation

This work is licensed under a Creative Commons Attribution 4.0 International License.