This dataset is a subset of MassiveSumm by subsampling and setting a maximum sequence length of 1.5k tokens. Links to reproduce the whole set of MassiveSumm via Common Crawl and the Wayback Machine are provided in the repository of MassiveSumm.
Description: MassiveSumm is a very large-scale, highly multilingual news summarization dataset designed to support summarization research across a wide range of languages. It encompasses 92 diverse languages and 35… See the full description on the dataset page:
https://huggingface.co/datasets/MaLA-LM/MassiveSumm_short.