Views
No views yet

| Model | Base Model | Link |
|---|---|---|
| Skywork-Reward-V2-Llama-3.1-8B | meta-llama/Llama-3.1-8B-Instruct | 🤗 Hugging Face |
| Skywork-Reward-V2-Llama-3.1-8B-40M | meta-llama/Llama-3.1-8B-Instruct | 🤗 Hugging Face |
| Skywork-Reward-V2-Llama-3.2-1B | meta-llama/Llama-3.2-1B-Instruct | 🤗 Hugging Face |
| Skywork-Reward-V2-Llama-3.2-3B | meta-llama/Llama-3.2-3B-Instruct | 🤗 Hugging Face |
| Skywork-Reward-V2-Qwen3-0.6B | Qwen/Qwen3-0.6B | 🤗 Hugging Face |
| Skywork-Reward-V2-Qwen3-1.7B | Qwen/Qwen3-1.7B | 🤗 Hugging Face |
| Skywork-Reward-V2-Qwen3-4B | Qwen/Qwen3-4B | 🤗 Hugging Face |
| Skywork-Reward-V2-Qwen3-8B | Qwen/Qwen3-8B | 🤗 Hugging Face |
| Category | Model | RewardBench v1 | RewardBench v2 | PPE Preference | PPE Correctness | RMB | RM-Bench | JudgeBench | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| Bradley-Terry | Llama-3-OffsetBias-RM-8B | 89.0 | 64.8 | 59.2 | 64.1 | 57.8 | 71.3 | 63.5 | 67.1 |
| ArmoRM-Llama3-8B-v0.1 | 90.4 | 66.5 | 60.6 | 60.6 | 64.6 | 69.3 | 59.7 | 67.4 | |
| Internlm2-20b-reward | 90.2 | 56.3 | 61.0 | 63.0 | 62.9 | 68.3 | 64.3 | 66.6 | |
| Skywork-Reward-Llama-3.1-8B-v0.2 | 93.1 | 71.8 | 62.2 | 62.5 | 66.6 | 72.1 | 62.9 | 70.2 | |
| LDL-Reward-Gemma-2-27B-v0.1 | 95.0 | 72.5 | 62.4 | 63.9 | 67.9 | 71.1 | 64.2 | 71.0 | |
| Skywork-Reward-Gemma-2-27B-v0.2 | 94.3 | 75.3 | 63.6 | 61.9 | 69.4 | 70.0 | 66.5 | 71.6 | |
| INF-ORM-Llama3.1-70B | 95.1 | 76.5 | 64.2 | 64.4 | 70.5 | 73.8 | 70.2 | 73.5 | |
| Generative | GPT-4o | 86.7 | 64.9 | 67.7 | - | 73.8 | - | 59.8 | - |
| Claude-3.5-Sonnet | 84.2 | 64.7 | 67.3 | - | 70.6 | - | 64.8 | - | |
| DeepSeek-GRM-27B | 88.5 | - | 65.3 | 60.4 | 69.0 | - | - | - | |
| DeepSeek-GRM-27B (w/ MetaRM) | 90.4 | - | 67.2 | 63.2 | 70.3 | - | - | - | |
| RM-R1-Qwen-Instruct-32B | 92.9 | - | - | - | 73.0 | 79.1 | - | - | |
| RM-R1-DeepSeek-Distill-Qwen-32B | 90.9 | - | - | - | 69.8 | 83.9 | - | - | |
| EvalPlanner (Llama-3.1-70B) | 93.9 | - | - | - | - | 80.0 | 50.9 | - | |
| EvalPlanner (Llama-3.3-70B) | 93.8 | - | - | - | - | 82.1 | 56.6 | - | |
| J1-Llama-8B | 85.7 | - | 60.3 | 59.2 | - | 73.4 | 42.0 | - | |
| J1-Llama-8B (Maj@32) | - | - | 60.6 | 61.9 | - | - | - | - | |
| J1-Llama-70B | 93.3 | - | 66.3 | 72.9 | - | 82.7 | 60.0 | - | |
| J1-Llama-70B (Maj@32) | - | - | 67.0 | 73.7 | - | - | - | - | |
| Bradley-Terry | Skywork-Reward-V2-Qwen3-0.6B | 85.2 | 61.3 | 65.3 | 68.3 | 74.5 | 74.4 | 67.6 | 70.9 |
| Skywork-Reward-V2-Qwen3-1.7B | 90.3 | 68.3 | 67.6 | 70.5 | 78.1 | 78.7 | 72.9 | 75.2 | |
| Skywork-Reward-V2-Qwen3-4B | 93.4 | 75.5 | 69.5 | 74.7 | 80.6 | 81.6 | 69.3 | 77.8 | |
| Skywork-Reward-V2-Qwen3-8B | 93.7 | 78.2 | 70.6 | 75.1 | 81.2 | 82.6 | 73.4 | 79.3 | |
| Skywork-Reward-V2-Llama-3.2-1B | 89.9 | 64.3 | 66.6 | 67.4 | 76.7 | 76.4 | 65.0 | 72.3 | |
| Skywork-Reward-V2-Llama-3.2-3B | 93.0 | 74.7 | 69.1 | 72.1 | 80.5 | 81.1 | 69.2 | 77.1 | |
| Skywork-Reward-V2-Llama-3.1-8B | 96.4 | 84.1 | 77.3 | 83.4 | 86.4 | 92.8 | 80.0 | 85.8 | |
| Skywork-Reward-V2-Llama-3.1-8B-40M | 97.8 | 86.5 | 79.8 | 87.2 | 89.3 | 96.0 | 83.4 | 88.6 |
[!NOTE] Although Skywork-Reward-V2-Llama-3.1-8B-40M outperforms the original Skywork-Reward-V2-Llama-3.1-8B, we consider it an experimental variant. This model is trained on the complete set of 40 million preference pairs, with about one third of the chosen-rejected pairs flipped. We recommend using this model solely for research or non-production purposes.
transformers1import torch
2
3from transformers import AutoModelForSequenceClassification, AutoTokenizer
4
5# Load model and tokenizer
6device = "cuda:0"
7model_name = "Skywork/Skywork-Reward-V2-Llama-3.1-8B"
8rm = AutoModelForSequenceClassification.from_pretrained(
9 model_name,
10 torch_dtype=torch.bfloat16,
11 device_map=device,
12 attn_implementation="flash_attention_2",
13 num_labels=1,
14)
15tokenizer = AutoTokenizer.from_pretrained(model_name)
16
17prompt = "Jane has 12 apples. She gives 4 apples to her friend Mark, then buys 1 more apple, and finally splits all her apples equally among herself and her 2 siblings. How many apples does each person get?"
18response1 = "1. Jane starts with 12 apples and gives 4 to Mark. 12 - 4 = 8. Jane now has 8 apples.
192. Jane buys 1 more apple. 8 + 1 = 9. Jane now has 9 apples.
203. Jane splits the 9 apples equally among herself and her 2 siblings (3 people in total). 9 ÷ 3 = 3 apples each. Each person gets 3 apples."
21response2 = "1. Jane starts with 12 apples and gives 4 to Mark. 12 - 4 = 8. Jane now has 8 apples.
222. Jane buys 1 more apple. 8 + 1 = 9. Jane now has 9 apples.
233. Jane splits the 9 apples equally among her 2 siblings (2 people in total). 9 ÷ 2 = 4.5 apples each. Each person gets 4 apples."
24
25conv1 = [{"role": "user", "content": prompt}, {"role": "assistant", "content": response1}]
26conv2 = [{"role": "user", "content": prompt}, {"role": "assistant", "content": response2}]
27
28# Format and tokenize the conversations
29conv1_formatted = tokenizer.apply_chat_template(conv1, tokenize=False)
30conv2_formatted = tokenizer.apply_chat_template(conv2, tokenize=False)
31# These two lines remove the potential duplicate bos token
32if tokenizer.bos_token is not None and conv1_formatted.startswith(tokenizer.bos_token):
33 conv1_formatted = conv1_formatted[len(tokenizer.bos_token):]
34if tokenizer.bos_token is not None and conv2_formatted.startswith(tokenizer.bos_token):
35 conv2_formatted = conv2_formatted[len(tokenizer.bos_token):]
36conv1_tokenized = tokenizer(conv1_formatted, return_tensors="pt").to(device)
37conv2_tokenized = tokenizer(conv2_formatted, return_tensors="pt").to(device)
38
39# Get the reward scores
40with torch.no_grad():
41 score1 = rm(**conv1_tokenized).logits[0][0].item()
42 score2 = rm(**conv2_tokenized).logits[0][0].item()
43print(f"Score for response 1: {score1}")
44print(f"Score for response 2: {score2}")
45
46# Expected output:
47# Score for response 1: 23.0
48# Score for response 2: 3.59375pip install "sglang[all]>=0.4.7.post1"NUM_GPUS GPUs are available):1NUM_GPUS=8
2for (( i=0; i<NUM_GPUS; i++ )); do
3 echo "Starting server on port $((8000+i)) with GPU: $i"
4 CUDA_VISIBLE_DEVICES=$i python -m sglang.launch_server \
5 --model-path Skywork/Skywork-Reward-V2-Llama-3.1-8B \
6 --mem-fraction-static 0.9 \
7 --tp 1 \
8 --host 127.0.0.1 \
9 --port $((8000+i)) \
10 --context-length 16384 \
11 --is-embedding \
12 &
13donetransformers example above.1import requests
2from transformers import AutoTokenizer
3
4
5model_name_or_path = "Skywork/Skywork-Reward-V2-Llama-3.1-8B"
6base_urls = [f"http://127.0.0.1:{8000 + i}/classify" for i in range(8)]
7tokenizer = AutoTokenizer.from_pretrained(model_name_or_path)
8
9
10def process_convs(convs, base_url, tokenizer, model_name_or_path):
11 payload = {"model": model_name_or_path}
12 convs_formatted = []
13 for conv in convs:
14 conv = tokenizer.apply_chat_template(conv, tokenize=False)
15 if tokenizer.bos_token is not None and conv.startswith(tokenizer.bos_token):
16 conv = conv[len(tokenizer.bos_token):]
17 convs_formatted.append(conv)
18
19 payload.update({"text": convs_formatted})
20 rewards = []
21 try:
22 responses = requests.post(base_url, json=payload).json()
23 for response in responses:
24 rewards.append(response["embedding"][0])
25 assert len(rewards) == len(
26 convs
27 ), f"Expected {len(convs)} rewards, got {len(rewards)}"
28 return rewards
29 except Exception as e:
30 print(f"Error: {e}")
31 return [None] * len(convs)
32
33
34prompt = "Jane has 12 apples. She gives 4 apples to her friend Mark, then buys 1 more apple, and finally splits all her apples equally among herself and her 2 siblings. How many apples does each person get?"
35response1 = "1. Jane starts with 12 apples and gives 4 to Mark. 12 - 4 = 8. Jane now has 8 apples.
362. Jane buys 1 more apple. 8 + 1 = 9. Jane now has 9 apples.
373. Jane splits the 9 apples equally among herself and her 2 siblings (3 people in total). 9 ÷ 3 = 3 apples each. Each person gets 3 apples."
38response2 = "1. Jane starts with 12 apples and gives 4 to Mark. 12 - 4 = 8. Jane now has 8 apples.
392. Jane buys 1 more apple. 8 + 1 = 9. Jane now has 9 apples.
403. Jane splits the 9 apples equally among her 2 siblings (2 people in total). 9 ÷ 2 = 4.5 apples each. Each person gets 4 apples."
41
42conv1 = [{"role": "user", "content": prompt}, {"role": "assistant", "content": response1}]
43conv2 = [{"role": "user", "content": prompt}, {"role": "assistant", "content": response2}]
44
45rewards = process_convs([conv1, conv2], base_urls[0], tokenizer, model_name_or_path)
46print(f"Score for response 1: {rewards[0]}")
47print(f"Score for response 2: {rewards[1]}")
48
49# Expected output:
50# Score for response 1: 23.125
51# Score for response 2: 3.578125yuhao.liuu at kunlun-inc dot com and liang.zeng at kunlun-inc dot com.1@article{liu2025skywork,
2 title={Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy},
3 author = {Liu, Chris Yuhao and Zeng, Liang and Xiao, Yuzhen and He, Jujie and Liu, Jiacai and Wang, Chaojie and Yan, Rui and Shen, Wei and Zhang, Fuxiang and Xu, Jiacheng and Liu, Yang and Zhou, Yahui},
4 journal={arXiv preprint arXiv:2507.01352},
5 year={2025}
6}