PPO-C (PPO with Calibrated Reward Calculation) is an RLHF algorithm to mitigate verbalized overconfidence in RLHF-trained Large Language Models.
PPO-C adjusts standard reward model scores during PPO training. It maintains a running average of past reward scores as a dynamic threshold to
classify responses, and adjusts the reward scores based on model expressed verbalized confidence.
Please refer to our preprint (
Taming Overconfidence in LLMs: Reward Calibration in RLHF) and
repo for more details.