165,847,195 Vietnamese documents from 4 public corpora, 370.2 GB of Parquet, one schema
This dataset is the pinned public Vietnamese corpora as gao read them, every source put to one contract and one schema, before any cleaning.
What is it
What is in it
Where the text came from
How it is laid out
Reading it
What you can build with it
One row
The columns
What this repo is
What ships and what does not
Things to know before you use it
What this is… See the full description on the dataset page:
https://huggingface.co/datasets/open-index/vitco.