Three datasets for image-to-3D structure prediction (SFT data for Qwen3-VL-4B-Instruct):
v11canon — single Infinigen objects (chairs, tables, cabinets, etc.)
scene_v3 — multi-object Infinigen scenes
physx_v11canon — PartNet-Mobility articulated objects, rotated to v11 chirality
v11canon/
train.jsonl # 7045 rows
test_seen.jsonl # 393 rows (held-out objects in seen categories)… See the full description on the dataset page:
https://huggingface.co/datasets/wenjingbian/structure_data.