Views
No views yet
brysgo/gol-grpo-diversified-37679423grpo_training_data.jsonl: prompt, thought, action, reward, and observed state trajectories.rl_trajectories_with_cot.jsonl: environment-side trajectory log with chain-of-thought separated from action JSON.logs/collection_summary.json: aggregate data-collection metrics.logs/grpo_reward_100_fixed.jsonl: reward log from the 100-step GRPO run.logs/train_grpo_100_fixed.log: training log for the completed fixed GRPO run.gol-grpo-100-fixed/: trainer metadata/config. Large model weights are included only when uploaded with --include-model and present locally.unknownunknownunknownunknownunknownunknown