Synthetic multimodal instruction data. Prompts come from laion/imagenet_variations;
images are generated with FLUX.1-schnell, encoded to 32 SEED-2 tokens per image,
grounded with Florence-2, and turned into USER / ASSISTANT records.
49,334 records / 93,684 turns / 71,015 images. This is the full run of the v4
design that the 990-record pilot
(imagenet-variations-synth-pilot-v4)
validated, rebuilt from 50,000 source prompts across a 25-task GPU array.… See the full description on the dataset page:
https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-v4.