🧪 ChemVQA-2K: A Visual Question Answering Dataset for Molecular Understanding
📘 Overview
ChemVQA-2K is a novel Visual Question Answering (VQA) dataset designed to bridge chemistry and multimodal AI.
It contains approximately 2,000 high-resolution molecular images (512×512) generated from valid SMILES strings, accompanied by 10 structured Q&A pairs per molecule, resulting in ~20,000 image-question-answer triplets.
Each image represents a 2D chemical structure rendered using RDKit, while each… See the full description on the dataset page:
https://huggingface.co/datasets/chandrabhuma/ChemVQA-2K.