This benchmark comprises 500 random samples of CodeUltraFeedback dataset.
The benchmark includes responses of multiple closed-source LLMs that can be used as references when judging other LLMs using LLM-as-a-Judge:
Please refer to our GitHub repository to evaluate your own LLM on CODAL-Bench using… See the full description on the dataset page:
https://huggingface.co/datasets/coseal/codal-bench.