MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling

TL;DR AI
2 min readKey summary
Researchers introduced MetaEarth-MM, a unified foundation model for remote-sensing image generation and translation across five modalities.
The model first infers latent scene representations from available observations, then generates target modalities for paired joint generation and any-to-any translation.
It is trained on the large global EarthMM dataset with 2.8 million images, helping address the lack of complete paired Earth-observation data.
Experiments show strong generalization, suggesting a scalable approach for remote-sensing data generation and downstream analysis.
