ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model
This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText.
ChineseWebText
Dataset Overview
We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is assigned a… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText.