Views
No views yet
global_step_100/actor from the completed
AppGen Random-GRPO dense-reward v38 run on four NVIDIA B200 GPUs.score_only. A subsequent
bitwise audit compared all 750 tensors (8,767,123,696 elements) and confirmed
that this step-100 export is exactly equal to the pinned Model E base checkpoint.
The held-out validation success rate was also unchanged at every evaluation:
11/193 (5.6995%).uedgpo_q3_model_e_awlegacy_random_dense_v38_b200x4_s100_20260728100namhokaist/appgen-qwen3-vl-8b-sft-ngc-amex-avariant-E-ngc-lr2p5e7-1ep6cdf0aa413850771f9a6f4c4da38f53d9d060f1cgrpo_ued_train at 750fce45b6dfb8550c352be950582a1bbccd654d2e-78; train rows: 8; PPO minibatch: 64250.0; adaptive replay disabled100; checkpoint/evaluation interval: 10appgen_system_prompt.txt4,0418795391af87a58a1670a33ae3f3e568e2af46899bc9b93fdc795c5f1fbbedfafluca0621/appgen-training-data@e96ee5f9f53b394a23ca1ce1912cb9ad40720842luca0621/appgen-sft-data@c195ae15abd3d6aaa07f971b8d732e7a28fa8dbfsft_qwen3_UNIFIED.json, 2,768 rows0f6817d5c522d629437389aeec02d4d4f1fa88e7b10dd9770ad9d9569557f8bf| Category | Success / tasks | Rate |
|---|---|---|
| communication | 0 / 28 | 0.00% |
| finance | 1 / 13 | 7.69% |
| food and drink | 1 / 17 | 5.88% |
| health and fitness | 0 / 18 | 0.00% |
| maps and navigation | 2 / 21 | 9.52% |
| productivity | 1 / 22 | 4.55% |
| shopping | 1 / 11 | 9.09% |
| social | 1 / 15 | 6.67% |
| tools | 3 / 20 | 15.00% |
| travel and local | 1 / 28 | 3.57% |
| pooled | 11 / 193 | 5.70% |
0.3516881456426913.appgen_grpo_manifest.json: exact run, prompt, base-model, and training contractgrpo_publication_identity.json: identity of this GRPO publicationstep100_vs_base_weight_audit.json: zero-update and exact-weight-equality auditrun_manifest.json: inherited Model E SFT provenancepublication_identity.json and training_verification.json: inherited Model E publication records