We have generated a hybrid multi-modal dataset consisting of over 500,000 instances for multi-modal training (Stage-2 training in our paper). You can download our dataset from this 🤗 HF Link.
Process the image compression package with the following commands:
cat images.tar.part* > images.tar
tar -xvf images.tar
If you obtain the following… See the full description on the dataset page:
https://huggingface.co/datasets/JUNJIE99/VISTA_S2.