The work introduces MarineEval, the first large-scale benchmark specifically designed to evaluate the marine understanding capabilities of Vision-Language Models (VLMs). MarineEval contains 2,000 expert-verified image-based question–answer pairs across 7 task dimensions and 20 domain-specific capacity dimensions, emphasizing specialized marine knowledge, visual reasoning, and real-world complexity. Through comprehensive benchmarking of 17 state-of-the-art… See the full description on the dataset page:
https://huggingface.co/datasets/WongYukKwan/MarineEval.