Views
No views yet
LlamaForSequenceClassification architecture to evaluate the honesty of steps in a dialogue. It was trained as part of a multi-objective alignment setup to allow for controllable inference and improved human value alignment.| Training Loss | Epoch | Step | Validation Loss | Accuracy |
|---|---|---|---|---|
| 0.3012 | 0.4632 | 50 | 0.2989 | 0.886 |
| 0.2851 | 0.9265 | 100 | 0.2657 | 0.908 |
1@article{shen2025simultaneous,
2 title={Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards},
3 author={Shen, Yiran and Xia, Yu and Chang, Jonathan and Ammanabrolu, Prithviraj},
4 journal={arXiv preprint arXiv:2510.01167},
5 year={2025},
6 url={https://arxiv.org/abs/2510.01167}
7}