This dataset contains structured evaluation sets used to study and benchmark the confidence behavior of large language models (LLMs). The dataset covers both multiple-choice and open-ended formats across diverse domains (e.g., clinical, law), with responses generated by a range of LLMs.
GitHub Repository:
https://github.com/ledengary/ccps
The dataset is organized by task type… See the full description on the dataset page:
https://huggingface.co/datasets/ledengary/CCPS.