This is the dataset corresponding to our paper "Benchmarking Generation and Evaluation Capabilities of Large Language
Models for Instruction Controllable Summarization".
The dataset subset contains 100 human-written data examples by us.
Each example contains an article, a summary instruction, a LLM-generated summary, and a hybrid LLM-human summary.
This subset contains human evaluation results for the 100 examples in the dataset… See the full description on the dataset page:
https://huggingface.co/datasets/Salesforce/InstruSum.