A model trained using RENT: Reinforcement Learning via Entropy Minimization - an unsupervised RL method that requires no external rewards or ground-truth labels. See our
github repo and
paper for more info on how this model was trained.
This model was trained using the
aime dataset.
When evaluating this model and the base model on AIME (64 runs on each model), we achieve the following results:
(Note that we report the mean and stderr of the 64 scores the model achieves on AIME)
This checkpoint has not been trained, evaluated, or tested on any other dataset.