Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one?
This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides:
Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page:
https://huggingface.co/datasets/ZaidGhazal/world-models-eval.