The mini-vncc is a 777,777 unique web documents represents an intensively filtered collection of Vietnamese web content, meticulously extracted from ~6TB of Vietnamese text in all CommonCrawl archive from 2013 to 2023. It is specifically tailored for pretraining models on Markdown structured web content in Vietnamese.
dataset = load_dataset("nampdn-ai/mini-vncc")… See the full description on the dataset page:
https://huggingface.co/datasets/nampdn-ai/mini-vncc.