This is a base model that has had an experimental reward model RL training done over it for a subset of the Erebus dataset (creative writing).
Reward function files can be found here:
verifiers
This model was trained using my chunked pref reward model baseline:
pretrain-rm-baseline-7b