CRUXEval: Code Reasoning, Understanding, and Execution Evaluation
π Home Page β’
π» GitHub Repository β’
π Leaderboard β’
π Sample Explorer
CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation) is a benchmark of 800 Python functions and input-output pairs. The benchmark consists of two tasks, CRUXEval-I (input prediction) and CRUXEval-O (output prediction). The benchmark was constructed as follows: first, we use Code Llama 34B to generate a large set of⦠See the full description on the dataset page: https://huggingface.co/datasets/cruxeval-org/cruxeval.