Views
No views yet
physics RLSD_TR batch-size-32 run.mean@16. checkpoints/last/ contains the final checkpoint.| Dataset | Method | Base model | Train batch size | Best val mean@16 | Best checkpoint | Final val mean@16 | Final checkpoint |
|---|---|---|---|---|---|---|---|
| Physics / SciKnowEval physics | RLSD_TR | Qwen3-8B | 32 | 68.05% | 90 | 65.78% | 100 |

| step | val_mean16 | percent |
|---|---|---|
| 10 | 0.586718750000 | 58.67% |
| 20 | 0.617187500000 | 61.72% |
| 30 | 0.614062500000 | 61.41% |
| 40 | 0.617968750000 | 61.80% |
| 50 | 0.642187500000 | 64.22% |
| 60 | 0.645312500000 | 64.53% |
| 70 | 0.653125000000 | 65.31% |
| 80 | 0.666406250000 | 66.64% |
| 90 | 0.680468750000 | 68.05% |
| 100 | 0.657812500000 | 65.78% |
| Section | Parameter | Value | Source |
|---|---|---|---|
| Run identity | Base model | Qwen/Qwen3-8B | queue/script override |
| Run identity | Dataset | Physics / SciKnowEval physics | run_qwen3_generalization.sh |
| Run identity | Method | RLSD_TR | run_qwen3_generalization.sh |
| Run identity | Config | rlsd | run_qwen3_generalization.sh |
| Run identity | Experiment | qwen3gen-physics-RLSD_TR-Qwen-Qwen3-8B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8 | run_qwen3_generalization.sh |
| Run identity | W&B run | run-20260703_055133-5wz4kyyd | wandb |
| Data | Train file | datasets/sciknoweval/physics/train.parquet | script override |
| Data | Validation file | datasets/sciknoweval/physics/test.parquet | script override |
| Data | Train batch size | 32 | queue/script override |
| Data | Train max samples | 3200 | queue/script override |
| Schedule | Total training steps | 100 | queue/script override |
| Schedule | Validation before train | False | queue/script override |
| Schedule | Save frequency | 10 | queue/script override |
| Schedule | Validation frequency | 10 | queue/script override |
| Sequence | Max prompt length | 2048 | queue/script override |
| Sequence | Max response length | 8192 | queue/script override |
| Sequence | Max model length | 10240 | queue/script override |
| Rollout | Train rollout n | 8 | queue/script override |
| Rollout | Validation rollout n | 16 | queue/script override |
| Rollout | vLLM GPU memory utilization | 0.8 | queue/script override |
| Optimization | Learning rate | 1e-6 | RLSD_TR method override |
| Optimization | Weight decay | 0.01 | script override |
| PPO/GRPO | PPO mini batch size | 8 | queue/script override |
| PPO/GRPO | Normalize GRPO advantages by std | False | baseline_grpo.yaml / script override |
| Rollout correction | Importance sampling mode | token | script override |
| Rollout correction | IS threshold | 2.0 | script override |
| Checkpoint/Logging | Checkpoint root | checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD_TR-Qwen-Qwen3-8B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8 | script override |
| Checkpoint/Logging | Latest checkpointed iteration | 100 | latest_checkpointed_iteration.txt |
| Checkpoint/Logging | External actor archive | checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD_TR-Qwen-Qwen3-8B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8/_actor_archive | preserve_actor_checkpoints.py |
| Checkpoint/Logging | Logger | console, wandb | ppo_trainer.yaml |
| RLSD_TR | Policy loss mode | rlsd | method override |
| RLSD_TR | Teacher regularization | trust-region | method override |
| RLSD_TR | Trust-region mix / teacher update rate | 0.1 | queue/script override |
| RLSD_TR | Token reweight lambda | 0.5 | queue/script override |
| RLSD_TR | Token reweight eps_w | 0.2 | queue/script override |
| RLSD_TR | Token reweight decay steps | 0 | queue/script override |
| RLSD_TR | Fused kernels | False | method override |
results/validation_mean16.csvresults/training_scores.csvresults/hyperparameters.csvresults/training_score.pngresults/training_score.svgartifacts/output.logartifacts/queue.log1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3repo_id = "SeongryongJung/Qwen3-8B-Physics-RLSD-TR"
4tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
5model = AutoModelForCausalLM.from_pretrained(
6 repo_id,
7 torch_dtype="auto",
8 device_map="auto",
9 trust_remote_code=True,
10)checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD_TR-Qwen-Qwen3-8B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD_TR-Qwen-Qwen3-8B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8/global_step_90/actorcheckpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD_TR-Qwen-Qwen3-8B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8/global_step_100/actorrun-20260703_055133-5wz4kyydartifacts/queue.log