The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset.
The dataset contains 100 Million <image, text> pairs
Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 )
Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration.
The original website: wukong-dataset.github.io
Terms of Use -… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Noah-Wukong-100M.