We encourage you to "heart" this reward model & use it in your multi-objective RLHF pipelines!
We release a custom bias reward model that is highly compute-efficient (only 0.1B parameters) and avoids knowledge degradation, providing a plug-and-play resource that can be seamlessly integrated into complex, multi-objective RLHF pipelines without conflicting with other objectives or adding compute overhead. Thus, this reward model lowers the barriers to entry and enables more researchers to implement robust bias mitigation into their RLHF pipelines without any compute or capability trade-offs.