This is the filtered Vietnamese subset of the DataComp Large Pool.
The table below shows the processing steps applied to achieve this subset.
Filtered for Vietnamese text using the fasttext-language-identification model (score threshold > 0.7)
26,537,765
100%
Removed images with a smaller dimension below 200… See the full description on the dataset page:
https://huggingface.co/datasets/minhnguyent546/datacomp_large_vie.