This is the dataset used by the automatic sparse attention compression method MoA.
It enhances the calibration dataset by integrating long-range dependencies and model alignment.
MoA utilizes long-contextual datasets, which include question-answer pairs heavily dependent on long-range content.
The question-answer pairs are written by human in this dataset repository. Large language Models (LLMs) should… See the full description on the dataset page:
https://huggingface.co/datasets/nics-efc/MoA_Long_HumanQA.