JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
Overview
JudgeSense is a benchmark dataset of 500 hand-validated prompt pairs for measuring prompt sensitivity in LLM-as-a-Judge evaluation systems. Each pair contains two differently phrased but semantically equivalent judge prompts applied to the same response, enabling rigorous measurement of how much a judge's decision changes due to prompt wording alone.
All 500 pairs were independently… See the full description on the dataset page: https://huggingface.co/datasets/anonymousreview111/judgesense-benchmark.