Validation data from the paper "Multi-Agent Judging for LLM Evaluation: A Data-Centric Analysis of Concordance with Human Preferences" by Anthony Boisbouvier.
This dataset contains 360 pairwise LLM evaluations judged by a panel of three frontier-class LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Flash) under blind conditions with Borda count aggregation, compared against human preference labels from MT-Bench and Chatbot Arena.… See the full description on the dataset page:
https://huggingface.co/datasets/anthonyboisbouvier-paris/agent-clash-multi-judge-eval.