Reformatted and deduped (e.g., alpaca13b vs gpt4 may have the same answer pair as alpaca13b vs gpt-3.5-turbo for some questions) from lmsys/mt_bench_human_judgments. This can be used as a quick evaluation metric to "measure" MT-bench performance during training.
Note the split names are converted to train_ and test_. Although the train_ splits will NOT be used to train anything, this split name makes some data… See the full description on the dataset page:
https://huggingface.co/datasets/when2rl/mt_bench_human_judgments_reformatted.