A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.