This model is a fine-tuned version of
Qwen/Qwen2.5-Math-7B-Instruct on the
mlfoundations-dev/putnam_bench_r1 dataset.
It has been trained using
TRL.
1from transformers import pipeline
2
3text = "The capital of France is Paris."
4rewarder = pipeline(model="dkjo8/Tars", device="cuda")
5output = rewarder(text)[0]
6print(output["score"])
This model was trained with Reward.
1@software{vonwerra2020trl,
2 title = {{TRL: Transformers Reinforcement Learning}},
3 author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
4 license = {Apache-2.0},
5 url = {https://github.com/huggingface/trl},
6 year = {2020}
7}