A Greek abstractive summarization dataset collected from the Greek part of Wikipedia, which contains 93,432 articles, their titles and summaries.
This dataset has been used to train our best-performing model GreekWiki-umt5-base as part of our research paper:Giarelis, N., Mastrokostas, C., & Karacapilidis, N. (2024) Greek Wikipedia: A Study on Abstractive Summarization.For information about dataset creation, limitations etc. see the original article.… See the full description on the dataset page:
https://huggingface.co/datasets/IMISLab/GreekWikipedia.