VLM Sample Dataset Card
Dataset details
train.json contains the multimodal synthesized conversation from the image-caption pairs, by adding randomly selected instructions
Primary intended uses:
The primary use of LLaVA is debug/test training framework for Chinese VLM models, like Qwen2VL.