A dataset of quality scores given by LLM judges and reward models for the outputs of many systems over the Arena Hard v0.1 benchmark.
These judgment scores are the raw data we collected for the JuStRank paper from ACL 2025, on system-level LLM judge performance and behavior, and also used to create the JuStRank Leaderboard.
In our research we tested 10 LLM judges and 8 reward models, and asked them to score the responses of 63 systems (generative… See the full description on the dataset page:
https://huggingface.co/datasets/ibm-research/justrank_judge_scores.