Views
No views yet
\boxed{N} where N is the option number.\boxed{N} matches ground truth)\boxed{N} exists but wrong answer)\boxed{} found)| Parameter | Value |
|---|---|
| base_model | Qwen/Qwen3-4B |
| algorithm | Async GRPO (TRL AsyncGRPOTrainer) |
| thinking | enabled (enable_thinking=True) |
| learning_rate | 3e-6 |
| lr_scheduler | cosine |
| warmup_steps | 60 |
| max_steps | 2000 |
| global_step_saved | 2000 |
| per_device_train_batch_size | 1 |
| gradient_accumulation_steps | 32 |
| num_train_processes | 4 |
| effective_batch_size | 128 prompts/step |
| num_generations | 9 |
| max_completion_length | 4096 |
| temperature | 1.0 |
| epsilon | 0.2 |
| epsilon_high | 0.2 |
| max_staleness | 4 |
| weight_sync_steps | 1 |
| max_grad_norm | 1.0 |
| precision | bf16 |
| parallelism | FSDP2 (4 GPUs training) + vLLM TP=4 (4 GPUs inference) |
| final_reward | ~0.45 |
| final_mean_completion_length | ~2370 tokens |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model = AutoModelForCausalLM.from_pretrained("EnergyAI/qwen3-4b-agrpo-think-lr3e-6", torch_dtype="auto")
4tokenizer = AutoTokenizer.from_pretrained("EnergyAI/qwen3-4b-agrpo-think-lr3e-6")