We provide 1M high-quality triplets of the form (flawed image, high-quality image, reflection) collected across
multiple domains using our scalable pipeline from [1]. We used this dataset to train our reflection tuning model.
To know the details of the dataset creation pipeline, please refer to Section 3.2 of [1].
Project Page:
https://diffusion-cot.github.io/reflection2perfection
We provide the dataset in the webdataset format for fast… See the full description on the dataset page:
https://huggingface.co/datasets/diffusion-cot/GenRef-wds.