NOVEReason is the dataset used in the paper NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning. It is a multi-domain, multi-task, general-purpose reasoning dataset, comprising seven curated datasets across four subfields: general reasoning, creative writing, social intelligence, and multilingual understanding. The data has been carefully cleaned and filtered to ensure suitability for training large reasoning models using… See the full description on the dataset page:
https://huggingface.co/datasets/thinkwee/NOVEReason_2k.