MegaWika 2 is an improved multilingual text dataset containing a structured view of Wikipedia articles, the web sources they cite, source text quality estimates, article text translations, and additional article enrichments.
Note: Web citations (sources) in the HuggingFace dataset do not include scraped source text; use rehydrate-citations.py to rehydrate them.
The initial data release is based on Wikipedia dumps from May 1, 2024.
In total, the data contains about 77… See the full description on the dataset page:
https://huggingface.co/datasets/jhu-clsp/megawika-2.