Views
No views yet
uedgpo_q3_model_e_simplepath_sparse_s10warm_v42.mean_std_norm, step-advantage weight 0), γ=1.0, group size 8,
KL loss low_var_kl coef 0.05, entropy 0, lr 5e-7, 1 ppo epoch.MobileGUI, sparse A2B reward (dense_reward=False),
simple-path curriculum (4 → 20 hops, +2 every 10 steps). PLR replay and ACCEL mutations are disabled
(replay_prob=0, no_adaptive_replay, no_mutations) — i.e. this is not a full PLR+ACCEL UED run.Qwen/Qwen3-VL-8B-Instructnamhokaist/appgen-qwen3-vl-8b-sft-ngc-amex-avariant-E-ngc-lr2p5e7-1ep (ckpt-113)namhokaist/appgen-qwen3-model-e-simplepath-sparse-h200x4-20260730-step10sft_exact prompt profile and Qwen3-VL mobile_use tool-call format with normalized 0–999 coordinates.
See appgen_system_prompt.txt in this repo.