This model is a fine-tuned version of
HuggingFaceTB/SmolLM-135M-Instruct on the
HumanLLMs/Human-Like-DPO-Dataset dataset.
It has been trained using
TRL.
1from transformers import pipeline
2
3text = "The capital of France is Paris."
4rewarder = pipeline(model="SupNek/trainer_output", device="cuda")
5output = rewarder(text)[0]
6print(output["score"])
This model was trained with Reward.
1@software{vonwerra2020trl,
2 title = {{TRL: Transformers Reinforcement Learning}},
3 author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
4 license = {Apache-2.0},
5 url = {https://github.com/huggingface/trl},
6 year = {2020}
7}