A benchmark for measuring whether an LLM judge returns the same verdict when
the same evaluation request is worded differently.
The question is deliberately narrow. This dataset does not measure whether a
judge is correct; it measures whether it is reproducible. A judge that is
wrong on every item but wrong identically under both phrasings scores perfectly
here, and that is intended — a measuring instrument whose reading depends on how
you phrase the question is not… See the full description on the dataset page:
https://huggingface.co/datasets/Rohithreddybc/judgesense-benchmark.