Format-standardized (v4) release of the synth_llava / synth_llava2
(image+caption) training text -- 603,999 rows combined. <seed2_N> payload
tokens are byte-identical to the prior release; only the surrounding text
framing changed.
(v3) This was the only one of the 6 sources with NO chat marker at all
previously (a bare
... Describe this image. instruction
with no USER:/ASSISTANT: framing). That release added USER:… See the full description on the dataset page:
https://huggingface.co/datasets/EmpathicRobotics/synth-llava.