ru-alpaca-eval is translated version of alpaca_eval. The translation of the original dataset was done manually. In addition, content of each task in dataset was reviewed, the correctness of the task statement and compliance with moral and ethical standards were assessed. Thus, this dataset allows you to evaluate the abilities of language models to support the Russian language. Baseline responses updated with GPT-4o model and also reviewed.
Overview of the… See the full description on the dataset page: https://huggingface.co/datasets/t-tech/ru-alpaca-eval.