The MT-Bench dataset is a collection of challenging multi-turn, open-ended questions designed to evaluate chat assistants and language models. Using LLM-as-a-judge, this dataset leverages strong models like GPT-4 to assess response quality and provide automated grading. This README provides details on using and extending the dataset for evaluation purposes.
There has been a proliferation of LLM-based chat assistants (chatbots) that leverage… See the full description on the dataset page:
https://huggingface.co/datasets/ZoneTwelve/mt-bench-tw.