CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding Capabilities
š CodeMMLU
CodeMMLU is a comprehensive benchmark designed to evaluate the capabilities of large language models (LLMs) in coding and software knowledge.
It builds upon the structure of multiple-choice question answering (MCQA) to cover a wide range of programming tasks and domains, including code generation, defect detection, software engineering principles, and much more.
šā¦ See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/CodeMMLU.