Views
No views yet
countdown/
4_30_100/oc_0.0/ GRPO on countdown-4-30-100, overconf coefficient 0.0
4_30_100/oc_1.0/ GRPO on countdown-4-30-100, overconf coefficient 1.0
4_10_50/oc_0.0/ GRPO on countdown-4-10-50, overconf coefficient 0.0
4_10_50/oc_1.0/ GRPO on countdown-4-10-50, overconf coefficient 1.0
math/
oc_0.0/ GRPO on dapo_hard, overconf coefficient 0.0
oc_1.0/ GRPO on dapo_hard, overconf coefficient 1.0oc_X.Y denotes the value of probe.overconf_coeff used during training.safetensors format and can be loaded directly with
AutoModelForCausalLM.from_pretrained(...).model_world_size_N_rank_*.pt). Use the verl checkpoint merger
(scripts/model_merger.py in the source repo) to consolidate them into a
single HuggingFace-loadable model.