Curated multimodal data for training a Qwen2.5-VL self-reflection RL pipeline,
plus the held-out LIVR splits and three external benchmarks (BLINK +
PixMo-Count + VSP) used to evaluate it.
train/ # LIVR train (9 tasks × 1000)
metadata.jsonl
livr_v2_manifest.json
images//... # ~8.7 GB
livr_eval/ # LIVR's own held-out val + test (8 tasks; counting → pixmo_count_eval)
validation/… See the full description on the dataset page:
https://huggingface.co/datasets/Kkuntal990/LIVR_mixed.