Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout:
per-step imagined frame (MSE target)
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page:
https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg.