As shown in the figure below, current video benchmarks often lack comprehensive annotations of CoT steps, focusing only on the accuracy of final answers during model evaluation while neglecting the quality of the reasoning process. This evaluation approach makes it difficult to comprehensively evaluate modelβsβ¦ See the full description on the dataset page:
https://huggingface.co/datasets/VLM-Reasoning/VCR-Bench.