This is a gradually self-truthified model (with 9 iterations) proposed in the paper
GRATH: Gradual Self-Truthifying for Large Language Models.
Note: This model is applied with DPO ten times. The reference model of DPO is set as the pretrained base model to avoid the overfitting problem.