Custom pseudo "fill in the middle" trained model, designed to handle varying "corruption rates" (randomized UTF8 character substitution).
Two custom GRPO reward functions were used to improve the pre-existing SFT trained model in order to have it more reliably attend to the XML styling.
The primary utility of this model is as a means to synthesize rejected / lower quality preference data from pre-existing SFT data (i.e, the general pretraining corpus).
This is useful in the context of teaching a reward model generalized preferences from lower quality, subtly incoherent base model-esque completions, of which are trivial to produce compared to human annotations.
Acknowledgements
Trained on 8xH200s provided free of charge by Deepshard for research & open source experimentation. Big McThankies.