Source Dataset: Derived from R1-Vision-Reasoning-Instructions with rigorous quality filtering.
Key Processing Steps:
Removed all examples containing bounding box annotations
Excluded pure image captioning tasks (difficult to verify)
Answers exceeding 5 words underwent model-based verification
Implemented special handling for LaTeX-formatted responses
Verified answer consistency using sequence… See the full description on the dataset page:
https://huggingface.co/datasets/le723z/Vision-Reasoning-QA.