MODUS is a large-scale, pixel-aligned 15-modality dataset for any-to-any
multimodal training. Every sample aligns 15 modalities covering appearance,
geometry, structure, segmentation, detection, text, and learned features.
Paper:
https://huggingface.co/papers/2607.25948
Code:
https://github.com/EPFL-VILAB/Modus
Structure
canny, sam_edge… See the full description on the dataset page:
https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality.