BERTJudge-Dataset is the training dataset used for developing BERTJudge models, as introduced in the paper BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation. It comprises question–candidate–reference triplets generated by 36 recent open-weight, instruction-tuned models across 7 established tasks, and synthetically annotated using nvidia/Llama-3_3-Nemotron-Super-49B-v1_5.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/BERTJudge-Dataset.