SceneChain-12K is a multi-turn scene editing conversation dataset for training vision-language models to generate and iteratively refine 3D indoor scenes.
messages: Multi-turn conversation following OpenAI chat format (system/user/assistant)
images: List of rendered scene image paths (relative to dataset root)
System: Scene editing instructions and tool definitions
User:… See the full description on the dataset page:
https://huggingface.co/datasets/runder1/SceneChain-12K.