Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Splits
N.o examples
Size
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB