Russian HumanEval (ruHumanEval) is the Russian analogue of the original HumanEval dataset, created to evaluate the ability of language models to generate code in the Python programming language to solve simple problems.
The dataset contains 164 tasks and is aimed at measuring the functional correctness of code generation based on information from the function's documentation lines — a text description of the function's operation and several… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/ruHumanEval.