750 tar shards, 123.5 GB total. One tar per scene id.
The source tree holds hundreds of thousands to millions of small files, far past the
100k files/repo the Hub recommends, so it is published as per-scene tar shards.
manifest.json scene id -> shard path and size
unpack.py restores the original directory tree
shards/.tar extracts to <scene_id>/<object_uuid>/
huggingface-cli… See the full description on the dataset page:
https://huggingface.co/datasets/nightgoodl/marco-aligned-v2.