Dataset for pre-training the LLM Quirel.
The code for processing data will be open-sourced later.
baidu baike data in Chinese, with Markdown structure
Collected from
https://huggingface.co/datasets/lars1234/baidu-baike-dataset
Cleaning procedure:
Organize context in Markdown.
Delete the paragraph with many network links.
Filter out texts that are too short.
wikipedia data in Chinese
Collected from… See the full description on the dataset page:
https://huggingface.co/datasets/RadiCat/wiki_pretrain.