š [ICLR '25] RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
This dataset contains the queries, generated checklist data, and responses data from 4 public benchmark datasets:
Dataset
No. of Queries
Comments
WildBench
1,000
To fit the context window of lightweight LLMs, we use a subset of WildBench including 1000⦠See the full description on the dataset page:
https://huggingface.co/datasets/wjkim9653/RocketEval-sLLMs.