The dataset used to train sft_4b_canon_aug2_v13mv_v13aug2_scenev4_mv —
a Qwen3-VL-4B fine-tune that predicts part-level kinematic structure
(materials, parent tree, joints, relations) from a single rendered view +
its colour segmentation map.
jsonl/ structure-prediction JSONLs
physx_canon_structure_train.jsonl 4 895 PhysX-Anything (canonical Z-up… See the full description on the dataset page:
https://huggingface.co/datasets/wenjingbian/data_5mix.