Views
No views yet
sft-5ep-offline-balanced-rl_5e-6-answer_onlylihaoxin2020/qwen3-4B-instruct-refiner-sftanswer_onlyopen_instruct/grpo_fast_refiner_sft.py| Parameter | Value |
|---|---|
| Learning rate | 5e-6 |
| LR scheduler | constant |
| Beta (KL penalty) | 0.001 |
| KL estimator | kl3 |
| Advantage normalization | standard |
| Samples per prompt (rollout) | 8 |
| Unique prompts per rollout | 32 |
| Mini batches | 1 |
| Epochs per batch | 1 |
| Per-device train batch size | 1 |
| Temperature | 1.0 |
| Seed | 42 |
| Async mode | true |
| Adam offload | true |
| vLLM sync backend | nccl |
| Parameter | Value |
|---|---|
| Max token length | 8192 |
| Max prompt token length | 6144 |
| Response length | 1024 |
| Pack length | 8192 |
| Parameter | Value |
|---|---|
| Verification reward | 10.0 |
| Non-stop penalty | false |
| Gate judge score with format bonus | false |
| Apply paper citation reward | true |
| Paper citation weight | 0.5 |
lihaoxin2020/refiner_rl (split: train)lihaoxin2020/refiner_rl (16 samples, split: test)Qwen/Qwen3.5-35B-A3B (via vLLM)