Views
No views yet
experiment_name=grpo_video_4node_full_v3_24f100k_8b_base_ppexplore_v1, global_step_140 (the keeper).0.4890 @ step 140 — statistically tied with the cold-start base (0.4918) and SFT-770 warmstart (0.4845). The token-dropout exploration that won +4.6 pt on OMR delivers no gain on video; this run sits in the same ~0.485 dead-heat band.vero compute_score. mean = macro-mean of the 3 bench accuracies.global_step_140 (peak):| metric | mean | videomme | holmes | perceptioncomp |
|---|---|---|---|---|
| ppexplore τ0.95 @140 (keeper) | 0.4890 | 0.656 | 0.465 | 0.346 |
| cold-base RL keeper @80 (sibling) | 0.4918 | 0.658 | 0.474 | 0.343 |
| SFT-770 RL keeper @180 (sibling) | 0.4845 | 0.657 | 0.451 | 0.345 |
| stock Qwen3-VL-8B ckpt-0 (zero-shot) | 0.4444 | 0.6426 | 0.4143 | 0.2762 |
Qwen/Qwen3-VL-8B-Instruct (cold-start, resume_mode=auto).ngquangtrung57/verl@videorl-mods. Fully-async GRPO: FSDP2 trainer + vLLM rollouter.score = 0.8·accuracy + 0.2·format (FORMAT_WEIGHT=0.2, FORMAT_MIN_THINK_CHARS=100). No KL penalty.ppo_mini_batch_size=16 × require_batches=4 × rollout.n=8 = 512 trajectories/step.GROUP_VIDEO_TRAIN_MC_24F100K (5 video-MC parquets, 24 frames / 100k pixels).1e-6, warmup 25 steps; total_epochs=2; clip_ratio 0.2 / 0.3 (clip_c=10.0); max_prompt_length=17408, max_response_length=16384; enforce_eager=true; gpu_memory_utilization=0.75; staleness 0.4.| key | value |
|---|---|
enable | true |
trigger_mode | high |
top_prob_threshold (τ) | 0.95 |
k_explore | 4 (of n=8 rollouts explore; 4 stay clean) |
prompt_exploration_prob | 0.5 |
deterministic | true |
perturb_prob | 1.0 |
mask_from_loss | true |
drop_top_k | 1 |
restrict_to_think_region | true |
selection_seed | 42 |
test_freq=10000); video val is offline full-set eval on a dedicated 8×H100 node (vLLM TP1), every 20 fit-steps.verl_fully_async (entity quangtrung5705-nanyang-technological-university-singapore). Train metrics only — video val is offline, not on W&B:
https://wandb.ai/quangtrung5705-nanyang-technological-university-singapore/verl_fully_async/runs/88lmybpkvideo-8b-grpo-base (0.4918) if you just want the best keeper. Multiple-choice video QA, <think>…</think> then-answer format. No safety/RLHF alignment beyond the base.1from transformers import AutoModelForImageTextToText, AutoProcessor
2
3model_id = "ngqtrung/video-8b-grpo-ppexplore"
4model = AutoModelForImageTextToText.from_pretrained(model_id, dtype="auto", device_map="auto")
5processor = AutoProcessor.from_pretrained(model_id)
6
7messages = [{
8 "role": "user",
9 "content": [
10 {"type": "video", "video": "clip.mp4"},
11 {"type": "text", "text": "Answer the multiple-choice question. Reason inside <think>...</think>, then give the final letter."},
12 ],
13}]
14inputs = processor.apply_chat_template(
15 messages, add_generation_prompt=True, tokenize=True,
16 return_dict=True, return_tensors="pt"
17).to(model.device)
18out = model.generate(**inputs, max_new_tokens=1024)
19print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])ngquangtrung57/verl@videorl-mods; fully-async GRPO (FSDP2 + vLLM).docs/experiments_summary_8b.md).