Austin Xu*, Srijan Bansal*, Yifei Ming, Semih Yavuz, Shafiq Joty (* = co-lead, equal contribution)
TL;DR: ContextualJudgeBench is a pairwise benchmark with 2,000 samples for evaluating LLM-as-judge models in two contextual settings: Contextual QA and summarization. We propose a pairwise evaluation hierarchy and generate splits for our proposed hierarchy.
To run evaluation on… See the full description on the dataset page:
https://huggingface.co/datasets/Salesforce/ContextualJudgeBench.