Résumé
The increasing availability of multi-modal data in Earth observation, such as optical and radar images, has enabled more comprehensive analyses for applications like land-cover classification. Despite the advantages, missing modality data can arise during inference due to factors like sensor malfunctions, cloud cover, or limited spatio-temporal coverage, presenting challenges to multi-modal model design. These issues can hinder model performance in the predictive tasks and its broad applicability. In this work, we propose a multi-modal co-learning framework designed to enhance single-modality predictions through multi-modal co-training, focused on the case where only partial modality data is available during inference. Our approach disentangles modality-invariant and modality-specific information by explicitly modeling shared and specific spaces using modality-dedicated encoders. To achieve this, we adopt a multi-task co-learning strategy with multiple loss functions promoting both discriminative and disentangled features. We evaluate our model using SPOT very high spatial resolution imagery and Sentinel-2 time series data for land-cover classification over a study site featured by contrasted landscape, located in the Indian Ocean, namely Reunion Island. The results demonstrate the effectiveness of our method with consistent improving over single-modality prediction. This work highlights the potential of multi-modal co-learning in the field of remote sensing, advancing land-cover classification, as it enables single-modality models to benefit from multi-modal data available at training time.