Synthetic vision-language instruction pilot from laion/imagenet_variations prompts.
Sample 1000 prompts from imagenet_variations
Add one style framing per prompt (website / textbook / diagram / …)
Generate image with FLUX.1-schnell
Encode with SEED-2 (<seed2_N>, 32 tokens)
Qwen2.5-VL-7B-Instruct looks at the actual image and writes instruction + response(modes: describe / QA / howto / reverse)
Format:… See the full description on the dataset page:
https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-pilot.