📄 Paper: Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
🧠 GitHub: AutoCaption
This repository provides the SFT training data and MCTS-VCB evaluation benchmark generated by the AutoCaption framework.
sft_data – for supervised fine-tuning of caption models
mcts_vcb – for evaluation using MCTS-generated captions and keypoints… See the full description on the dataset page:
https://huggingface.co/datasets/HasuerYu/AutoCaption.