HumanEvalComm: Benchmarking the Communication Skills of Code Generation for LLMs and LLM Agent
π Paper β’
π» GitHub Repository β’
π€ Dataset Viewer
Dataset Description
HumanEvalComm is a benchmark dataset for evaluating the communication skills of Large Language Models (LLMs) in code generation tasks. It is built upon the widely used HumanEval benchmark. HumanEvalComm contains 762 modified problem descriptions based on the 164 problems in the⦠See the full description on the dataset page: https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm.