Potential mislabeled benchmark items surfaced by the paper "Auditing LLM Benchmarks with Item Response Theory".
Paper:
https://arxiv.org/abs/2605.30504
Rows are included when either delta_li > 0 or the GPT-5.4 weak-reference label is mislabel or unsure.
This is the union of items flagged by the unsupervised indicator and items flagged by the weak-reference labeler.
For items flagged only by the weak-reference labeler but filtered out… See the full description on the dataset page:
https://huggingface.co/datasets/Writer/IRT-mislabeled-items.