Everything here comes from one project: can a Qwen3-VL-2B-Instruct model improve at
vision-language benchmarks by writing and answering its own questions about unlabeled images,
with no human questions and no human answers anywhere in the loop?
Contents let you (A) reproduce a reported score in ~30 minutes with no training, (B) rerun the
training that produced it, (C) run the other arms, and (D) regenerate… See the full description on the dataset page:
https://huggingface.co/datasets/ahmedheakl/random.