The TRM-Preference dataset is introduced in the paper Characterizing, Evaluating, and Optimizing Complex Reasoning.
The dataset is designed to evaluate and optimize the quality of reasoning traces in Large Reasoning Models (LRMs) by training a Thinking Reward Model (TRM). Instead of focusing solely on answer correctness, TRM-Preference uses the ME² principle to evaluate "how a model thinks" across four dimensions:
Macro-Efficiency: Disciplined global structure… See the full description on the dataset page:
https://huggingface.co/datasets/zzzhr97/TRM-Preference.