Practical Real-world Industry and Multi-domain Evaluation benchmark.
PrimeBench is a benchmark for evaluating evaluators. Each of its 400 examples is a pair of
responses to the same prompt, deliberately edited so that one is better than the other along a
named criterion. A reward model or LLM judge passes an example if it scores the chosen response
above the rejected one.
Built and maintained by Composo.
Most preference datasets score… See the full description on the dataset page:
https://huggingface.co/datasets/ComposoAI/PrimeBench.