Views
No views yet
oct-low-agreeableness-llama3.1-8b-dpo-sft-r4.| Model | Method | Description |
|---|---|---|
oct-low-agreeableness-llama3.1-8b-sft-r4 | SFT only | This model. Single-stage character training. |
oct-low-agreeableness-llama3.1-8b-dpo-sft-r4 | DPO + SFT | Full two-stage pipeline per the OCT paper. |
1@article{OCT2024,
2 title={Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI},
3 author={Maius Haiduc},
4 year={2024},
5 journal={arXiv preprint arXiv:2511.01689}
6}