The Korean version of the HumanEval benchmark, where the original docstrings are translated into Korean.
Used DeepL translator API for translation.
Performed manual post-processing and LLM-based quality evaluation after translation.
Data format and benchmark performance evaluation method are identical to those of HumanEval.
Related Paper: KR-HumanEval을 활용한 언어 모델의 한국어 프로그램 합성 성능 분석 (HCLT 2024, best paper award)
Benchmark Version… See the full description on the dataset page: https://huggingface.co/datasets/HYU-NLP/KR-HumanEval.