This dataset is used for RubricARROW SFT training as presented in the paper RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains.
This dataset contains SFT training data for the RubricARROW judge model. Each example is formatted in an instruction-tuning style.
To extract the unique instructions… See the full description on the dataset page:
https://huggingface.co/datasets/OpenRubrics/RubricARROW-Judge-SFT.