The MT-Bench-Hi (Hindi MT-Bench) dataset is a multi-turn question set containing 200 prompts in the Hindi language to evaluate the conversational ability of the Hindi large language models (LLMs). The dataset has 80% of samples created natively by specialists well-versed in Hindi and 20% of the samples that are translated from the English version of the dataset.
This dataset is ready for commercial/non-commercial use. The evaluation steps are described here.… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/MT-Bench-Hi.