DepthVLM-Bench is a unified indoor-outdoor metric depth estimation benchmark designed for vision-language models (VLMs). The benchmark provides diverse indoor and outdoor scenes with metric depth annotations in a unified VLM-compatible format, enabling large multimodal models to jointly learn dense geometry prediction and multimodal understanding.
Unified indoor and outdoor metric depth estimation
VLM-compatible data format
Dense depth⦠See the full description on the dataset page:
https://huggingface.co/datasets/JonnyYu828/DepthVLM-Bench.