A large-scale dataset created by combining educational mathematics problems and web content to create a continued pretraining dataset for LLMs. The dataset contains 2,000,000 high-quality examples sorted by content relevance scores.
Data Source
Math Content: Hugging Face finemath-4plus dataset (educational math problems)
Web Content: Hugging Face fineweb-edu dataset (web pages from Common Crawl)
Training… See the full description on the dataset page: https://huggingface.co/datasets/qingy2024/DraftWeb-v1.