This dataset is an enriched, cleaned, and metadata-enhanced version of zwq2018/Multi-modal-Self-instruct.It pairs images with natural language questions and answers, making it ideal for Vision-Language Model (VLM) training, benchmarking, and instruction tuning.