Views
No views yet
| Training Loss | Epoch | Step | Validation Loss | Accuracy |
|---|---|---|---|---|
| 0.5633 | 1.0 | 22660 | 0.6624 | 0.6857 |
@article{sun2024supervised,
title={Supervised Fine-Tuning as Inverse Reinforcement Learning},
author={Sun, Hao},
journal={arXiv preprint arXiv:2403.12017},
year={2024}
}