DiffFuSR is a modular pipeline designed for super-resolving all 12 spectral bands of Sentinel-2 Level-2A imagery to a unified ground sampling distance (GSD) of 2.5 meters. The pipeline operates in two stages:
A diffusion-based super-resolution (SR) model trained on high-resolution RGB imagery from the NAIP and WorldStrat datasets, harmonized to simulate Sentinel-2 characteristics.
A learned fusion network that upscales the remaining multispectral bands by utilizing the super-resolved RGB image as a spatial prior.
This approach introduces a robust degradation model and contrastive degradation encoder to support blind SR, outperforming current SOTA baselines in reflectance fidelity, spectral consistency, spatial alignment, and hallucination suppression.