RWML (arXiv:2602.05842) reproduction data from
Ch1nyzzz/e2-learning, collected with
Qwen/Qwen2.5-7B-Instruct rollouts on ALFWorld:
rwml_grpo/train.parquet, rwml_grpo/val.parquet — prompts for RWML GRPO
training (filtered merged-10k corpus);
rwml_tau_calibration.json — calibrated embedding-distance threshold tau_d;
rwml_alfworld_qwen25_7b_validation_merged10k.jsonl — validation transitions.
Paths mirror the source repository's data/ layout; pull them back… See the full description on the dataset page:
https://huggingface.co/datasets/erv1n/e2l-rwml-alfworld-data.