This set of evaluation prompts is created by the LMSYS org for better evaluation of chat models.
For more information, see the paper.
To load this dataset, use 🤗 datasets:
from datasets import load_dataset
data = load_dataset(HuggingFaceH4/mt_bench_prompts, split="train")
To create the dataset, we do the following for our internal tooling.
rename turns to prompts,
add empty reference to remaining prompts… See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceH4/mt_bench_prompts.