Backup of sampling_dataset from the video_gen_physics workspace. The tree has too many
small files to commit individually (the Hub allows 128 commits/hour/repo), so it is
stored as independent tar shards of ~8GB under
shards/. Files too large to bundle are stored verbatim under large/.
manifest.txt lists how many files each shard holds.
Paths inside every tar are relative to sampling_dataset, so extract all shards into one
directory that plays the role… See the full description on the dataset page:
https://huggingface.co/datasets/doanh25032004/vgp_sampling_dataset_tar.