This fine-tuned version of Qwen2-VL-2B integrates both text and image modalities to improve multimodal reasoning capabilities. The model was fine-tuned on the PubMed Vision dataset using 6,666 paired samples of textual and visual data, enhancing its understanding of medical and visual-language tasks.
🧩 Applications
Visual Question Answering (VQA)
Medical Image Interpretation
Text-to-Image Contextual Understanding
Multimodal Research Tasks
💖 Credits
This model was built and fine-tuned with love using Unsloth and Hugging Face tools.