This repo contains the dataset MMDuetIT, which is used for training MMDuet, and benchmarks for evaluating MMDuet. The data distribution of MMDuetIT is as follows:
Dense Captioning
Shot2Story: 36949 examples from human_anno subset
COIN: 4574 examples from the train set with 2-4 minutes videos