Dataset introduced in the paper: Analyzing the Dialect Diversity in Multi-document Summaries (COLING 2022)
Olubusayo Olabisi, Aaron Hudson, Antonie Jetter, Ameeta Agrawal
DivSumm is a novel dataset consisting of dialect-diverse tweets and human-written extractive and abstractive summaries. It consists of 90 tweets each on 25 topics in multiple English dialects (African-American, Hispanic and White), and two reference summaries per input.… See the full description on the dataset page:
https://huggingface.co/datasets/Bisi/DivSumm.