This is our model using our algorithm iw-SFT (importance weighted SFT). For more details of the algorithm please refer to our paper below.
At a high level, the algorithm is motivated by showing first that SFT and RL are connected: the SFT objective lower bounds the RL objective. We show that by reweighting the curated dataset adaptively during training, i.e. iw-SFT, we can obtain a much tighter bound to RL than SFT alone.
We benchmark our algorithm on maths reasoning tasks, see table below.
For all links and general information see [here](For more general information see
here.