This repository contains the training dataset for the paper: Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs.
Code:
https://github.com/Haochen-Wang409/Grasp-Any-Region
The Grasp Any Region (GAR) dataset is designed to empower Multimodal Large Language Models (MLLMs) with comprehensive region-level visual understanding. While MLLMs excel at holistic understanding, they often struggle with… See the full description on the dataset page:
https://huggingface.co/datasets/HaochenWang/Grasp-Any-Region-Dataset.