A uniformly mixed Chinese pretraining dataset with 20B tokens, compiled from Ultra-FineWeb zh (67% by tokens, 54% by rows) and Ultra-FineWeb-L3 zh (33% by token, 46% by row). Every shard contains the same proportion of web and L3 documents.
Since Ultra-FineWeb is simply filtered from its source datasets and has not been deduplicated, we performed global near-deduplication on its subset. Additionally, we removed contents with too many non-Chinese characters… See the full description on the dataset page:
https://huggingface.co/datasets/UndefinedCpp/ultrafineweb-mix-20b.