[π Homepage] | π€ Dataset | π€ Paper | π arXiv | π» GitHub
The evaluation code can be found in π» GitHub.
[Abstract]
As recent multi-modality large language models (MLLMs) have shown formidable proficiency on various complex tasks, there has been increasing attention on debating whether these models could eventually mirror human intelligence. However, existing benchmarks mainly focus on evaluating⦠See the full description on the dataset page:
https://huggingface.co/datasets/Songweii/M3GIA.