MTC-Bench is a multimodal evaluation dataset pack designed for benchmarking vision-language models (image) and video-language models (video) across general multimodal capabilities and specific capabilities (fine-grained vision, reasoning/math, spatial structure reservation, holistic coverage, high-rate compression, and complex cross-modal alignment).
This repository provides a ready-to-run folder layout and lmms-eval task configs (via the mtcbench/ task folder) to enable… See the full description on the dataset page:
https://huggingface.co/datasets/DuHeng0128/MTC-Bench.