PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout —
per-step imagined frame (MSE target) + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page:
https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_noncot_chunk_k5_20260622_perseg.