This dataset is a synthetic dataset created using the Unity engine, specifically designed for depth estimation tasks in dense crowd scenarios. It includes left and right images, as well as the ground truth depth maps for the left images. The simulated camera first captures the left image from the left position, then moves one meter horizontally to the right to capture the right image. The camera's focal length is 600 pixels.… See the full description on the dataset page: https://huggingface.co/datasets/jibingyang111/PersonDataset.