This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to
https://hplt-project.org/datasets/v2.0
The Cleaned variant of HPLT Datasets v2.0
This is the cleaned variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here.
The original JSONL files… See the full description on the dataset page:
https://huggingface.co/datasets/jobs-git/HPLT2.0_cleaned.