This repository hosts the official multilingual pretraining datasets for NeoBabel.
This dataset is part of the work presented in the paper:NeoBabel: A Multilingual Open Tower for Visual Generation.
🔥 Official multilingual pretraining dataset for NeoBabel.
The repository is organized by dataset. Each dataset has… See the full description on the dataset page:
https://huggingface.co/datasets/mderakhshani/NeoBabel-Pretrain.