MSVD-Video-Captioning-Vi is a Vietnamese video captioning dataset derived from the MSVD dataset originally hosted by the user friedrichor on Hugging Face.
This dataset provides Vietnamese captions for short video clips and is intended for:
Video captioning research
Vision–Language model training
Multimodal instruction tuning
Video-to-text generation