A small synthetic dataset for experimenting with transparent and reproducible evaluation of large language model responses.
The dataset contains prompt-response pairs with manually constructed quality annotations covering response quality, uncertainty, overconfidence, and common evaluation failure modes.
Dataset purpose
This dataset accompanies the LLM Evaluation Lab Hugging Face Space.
It is designed primarily for: