LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
π Home Page β’
π» GitHub Repository β’
π Leaderboard β’
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is⦠See the full description on the dataset page: https://huggingface.co/datasets/mathewmouchamel/code_generation_v2.