This dataset evaluates candidate answers for various question-answering (QA) tasks across multiple datasets such as Jeopardy!, hotpotQA, nq-open, narrativeQA, and BIOMRC, etc. See details in paper. It contains questions, reference answers (ground truth), model-generated candidate answers, and human judgments indicating whether the candidate answers are correct.
question
string
The question asked… See the full description on the dataset page:
https://huggingface.co/datasets/zli12321/pedants_qa_evaluation_bench.