Views
No views yet
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B on math, exported from a single training step for evaluation purposes./opt/dlami/nvme/workspace/shapley_workspace/shapley/checkpoints/shapley_baselines/drgrpo_noent_cl20_r1distill_qwen1.5b_math_ec2_run10_0430_1638/global_step_231run10_0430_1638train_drgrpo_noent_cl20_r1distill_qwen1.5b_math_ec2.shpython -m verl.model_merger merge --backend fsdp --local_dir <ckpt>/actordeepseek-ai/DeepSeek-R1-Distill-Qwen-1.5Bnoent)cl20)shengjia-toronto/shapley-experiments). This is a baseline checkpoint (Dr.GRPO), used for comparison against Shapley-shaped credit assignment.deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B).