A benchmark of 10,843 equivalence groups designed to measure how consistently large language models answer semantically equivalent questions phrased in different surface forms.
Each group contains 2–5 restatements of the same question with the same correct answer.
The benchmark spans 39 question categories across factual retrieval and logical reasoning.
The dataset is procedurally generated by scripts/generate_ss_groups.py (seed 42) and is fully… See the full description on the dataset page:
https://huggingface.co/datasets/Hravan/semantic-sensitivity.