π Paper | π©π»βπ» GitHub | π€ Huggingface Datasets | π Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends⦠See the full description on the dataset page:
https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-hard.