This dataset is introduced in LLMLingua-2 (Pan et al., 2024), and is collected to construct the training data for LLMLingua-2 compressor.
It consists of 5169 instances from MeetingBank training split, with their GPT-4 compressed versions.
Given pairs of original texts and their compressed versions, we release the data annotation tool here to assign a binary label to each token in the original texts to determine if it should be preserved or… See the full description on the dataset page:
https://huggingface.co/datasets/microsoft/MeetingBank-LLMCompressed.