2,279,914 Vietnamese documents from 2 public corpora, 10.6 GB of Parquet, one schema
This dataset is the same corpora after the cleaning line: normalized, measured, filtered to Vietnamese prose, deduplicated on identity, and with the personal identifiers covered.
What is it
What is in it
Where the text came from
How it is laid out
Reading it
What you can build with it
One row
The columns
What this repo is
What ships and what does not
Things… See the full description on the dataset page:
https://huggingface.co/datasets/open-index/vitco-clean.