Photorealistic synthetic dataset (rendered in NVIDIA Isaac Sim) used to train
X-Lens. Each scene is captured by a
6-camera rig — 4 fisheye + 2 pinhole — with metric ground-truth depth, so a single
sample already contains the pinhole / fisheye / heterogeneous mix the model targets.
To stay friendly to the Hub, each scene is a single .tar (contents are already
compressed JPG/PNG/NPY, so the tar is stored uncompressed):
train/
.tar… See the full description on the dataset page: https://huggingface.co/datasets/henryzhou998/OmniScene.