Task sets for RL / OPD / RL+OPD runs on the Slime agent_envs stack. Rows are
stored in the Slime-readable schema so train.py can load them directly with
--input-key prompt --label-key label --metadata-key metadata.
prompt (model input, raw text): the fixed instruction the model receives.
The live per-turn… See the full description on the dataset page:
https://huggingface.co/datasets/huzican/agent_envs.