This dataset contains precomputed Qwen2.5-VL ViT features for ScanNet RGB frames, used by Qwen-3D as presented in Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding to skip the ViT forward pass during training.
Project page:
https://qwen-3d.github.io/Code:
https://github.com/ll220/qwen3dModel:
https://huggingface.co/katefgroup/Qwen-3DPaper:
https://arxiv.org/abs/2608.02980
See docs/RUN.md in the code repository for full… See the full description on the dataset page:
https://huggingface.co/datasets/nislamsm/CasualBench.