To download all WARC data from Common Crawl then filter out Vietnamese in Markdown and Plaintext format.
There is 1% of Vietnamse in CC, extract all of them out should be a lot (~10TB of plaintext).
To make use of raw data from common crawl, you need to do filtering… See the full description on the dataset page:
https://huggingface.co/datasets/Symato/cc.