DHP Benchmark: Are LLMs Good NLG Evaluators? 2408.13704
We present this DHP benchmarking dataset to evaluate the capablities of LLMs as NLG evaluators. We will release the evaluation prompts and code soon.
This dataset includes 6 subsets, covering four NLG tasks: Summarization (SummEval, SumPubMed), Completion (Story Cloze), Question Answering (Answer Equivalence), and Translation (WMT22-zhen, WMT22-deen).
Each subset includes… See the full description on the dataset page:
https://huggingface.co/datasets/YCWANGVINCE/DHP_Benchmark.