MMTab is a large-scale multimodal instruction-tuning dataset for enhancing and evaluating the visual table understanding ability of multimodal LLMs. It requires the model to generate correct responses based on a table image and input requests.
More details about MMTab can be found in the ACL 2024 paper Multimodal Table Understanding.