Our data is mostly the same as EmbedLLM's. Specially, we applied a majority vote to consolidate multiple answers from a model to the same query.
If the number of 0s and 1s is the same, we prioritize 1.
This step ensures a unique ground truth for each model-query pair… See the full description on the dataset page:
https://huggingface.co/datasets/JianhaoNJU/IrtNet-Dataset.