The LLMEval²-Select dataset is a curated subset of the original LLMEval² dataset introduced by Zhang et al. (2023). The original LLMEval² dataset comprises 2,553 question-answering instances, each annotated with human preferences. Each instance consists of a question paired with two answers.
To construct LLMEval²-Select, Zeng et al. (2024) followed these steps: