A curated collection of 4.79 million Wikipedia articles spanning the 2008 and 2010 snapshot releases, cleaned and compressed for efficient large-scale language model pretraining. This dataset preserves the raw encyclopedic knowledge of two distinct eras of Wikipedia, making it valuable for temporal analysis, knowledge evolution research, and foundation model training.
Split
Articles
Compressed Size
Raw JSONL… See the full description on the dataset page:
https://huggingface.co/datasets/AdhyanshVerma/wiki-2008-2010.