RLVE × Qwen3-32B, K=1024 采样(SFT 数据)
Qwen3-32B 在 RLVE 训练集 9000 道题上采样, 每题 2 条, 每条上限 1024 token。
共 18000 条。
与本项目所有 RLVE 评测逐项一致, 以避免训练/评测的 prompt 错配:
prompt: tokenizer.apply_chat_template(msgs, add_generation_prompt=True),
不传 enable_thinking(即 Qwen3 默认开启 thinking)
temperature 0.7, top_p 0.9, top_k -1, n=2, max_tokens 1024
含
0 / 18000 (0.0%)
含 \boxed{}
0 / 18000 (0.0%)
response 字符数
中位 3818 / 均值 3708 / 最大… See the full description on the dataset page:
https://huggingface.co/datasets/SeanWang0027/rlve-32b-k1024-sft.