Original data from:
https://github.com/armancohan/long-summarization
The first 3000 rows of the test split of the original dataset were processed and filtered as follows.
In the original dataset, some sentences appear several times in the same article, even if they're only contained once in the original research paper.
For this reason, all dataset rows where the same sentence appeared more than once where removed.
In the original dataset, every sentence is a separate string, and these strings… See the full description on the dataset page:
https://huggingface.co/datasets/giuliadc/pubmed-filtered.